Technical Excellence

RTO and RPO:
What You Are Allowed to Lose

Part VII · The Numbers That Decide Production readiness · Recovery

At the top of my banking demo's compose file there is a comment saying that Redis has no volume, because all of its state can be rebuilt: sessions by logging in again, idempotency keys because the database's unique index is the source of truth, caches because they refill. Read as a recovery objective, that comment is complete. It names the store, says everything in it may be lost, and gives the reason. The Postgres service, a few lines further down, holds the accounts and the ledger, has a volume, and has no such sentence.

This page is about the distance between those two services, and about the two numbers that fill it: how much data a system may lose, and how long it may take to come back. The demo is the one behind Not Run Is Not Passing and Money Has Rules the Framework Does Not Know, and I read it again with one question: what would each part of it answer if the disk went?

// the crux

You may lose what you can rebuild. Everything else needs a date on which you restored it.

// in one breath
  • A comment in a compose file that answers the question for one store, and a service a few lines further down that was never asked.
  • The number everyone asks for first, zero, and what its fine print charges on every commit.
  • A migration where nothing could be lost, and the three habits that showed it.
the numbers

Two Numbers, and Whose They Are

NIST's contingency planning guide, SP 800-34, states the two definitions plainly. The recovery point objective is a point in time: how far back the data may be once the system is running again. The recovery time objective is a length of time: how long the system may stay in recovery before the mission suffers. One is a moment, and the other is a duration.

Neither belongs to the people who build the system. AWS's Reliability Pillar says every workload needs both, chosen from the business impact of downtime and of lost data, and it lists the usual ways to get this wrong: objectives chosen arbitrarily, unrealistic ones such as zero data loss, and objectives stricter than the business needs, which buys a costlier and more complicated recovery than the workload asked for. The worksheet it hands you is a list of questions about what an outage and a lost stretch of data would cost the business, asked with the technical team in the room to say what is achievable.

My standards, which I wrote up with an AI assistant, carry the item in one line: “Backup and restore, with RPO and RTO.” It sits under the heading that means design it and write it down, and build it only if I ask.

One of those AWS questions is the one my compose comment answered without naming it: can the lost data be recreated from somewhere else? AWS calls that derived data and says to look at the recovery point of whatever it is derived from. Sessions come back when people log in again. A cache refills. An idempotency key comes back from the database, and the code beside the guard says why: Redis is only the fast path, and the unique index in Postgres is the source of truth, so a Redis flush cannot double-post. Those three kinds of state can be given the loosest objective there is, everything, and it costs nothing to meet.

Postgres holds the accounts, the transfers and the ledger, and nothing else in the stack can recreate them. That is where the two numbers start to cost money.

the price

What Each Way of Keeping Data Promises

Every way of keeping a copy has a window in which a failure costs you data, and the documentation says how wide the window is. It is different for each way.

Five ways of keeping a copy, and what a failure can cost under each
How the copy is keptWhat a failure can costWhere that is written
A nightly backupUp to a day of writes, because the newest recovery point is the last backupAWS, disaster recovery options: how often you back up sets your achievable recovery point
WAL archivingWhatever sits in a segment not yet archived, a long delay on a quiet system unless archive_timeout caps itPostgreSQL, continuous archiving: about a minute is called reasonable
Asynchronous commitUp to three times the WAL writer delay, so 600 milliseconds at the default of 200, of commits already reported donePostgreSQL, asynchronous commit
A synchronous standbyNo acknowledged commit, unless the primary and every synchronous standby lose their storagePostgreSQL, the synchronous_commit setting
A second Region, activeClose to nothing when a Region is lost; a corrupted or deleted row is copied along with everything elseAWS, disaster recovery options

PostgreSQL is careful to say that the third row risks data loss and not data corruption: after a crash the database is consistent and simply missing the last stretch of commits.

Recovery time runs the other way round from cost. AWS orders its four strategies from cheap and slow to expensive and fast: restore from backups, keep a pilot light on, run a warm standby, run active in two Regions. Under the first, the recovery time is however long a restore and a redeploy take, which nobody knows until someone has done one.

My demo is a single Postgres container on the default commit setting, which is none of the five rows. Postgres does not report a commit until it has flushed the log to disk, so a crash of the process loses nothing that was acknowledged. The performance notes price that flush at roughly one to five milliseconds a commit on the virtual disk Docker Desktop provides, an estimate in the notes and not a measurement. The disk itself has no such answer. The data lives on one named Docker volume, and the stack keeps no second copy of it.

the fine print

The Fine Print on Zero

Zero is the number everyone asks for first. AWS gives it as an example of an objective that may not be achievable, and its disaster recovery paper says why for data disasters: a copy that is kept current copies mistakes too. When a row is corrupted or deleted, the replica soon agrees with the mistake, and the last good recovery point is one from before anyone noticed. For failover in general, the same paper says that even with the best practice it describes, both numbers stay above zero.

The price arrives on every commit. Waiting for a standby to confirm each one puts a network round trip inside every write; waiting for the local disk puts a flush inside every write. My demo pays the second, at that estimated one to five milliseconds, and has no second machine to pay the first with.

A zero that holds up has three parts: the failure it covers, the price it charges per commit, and the date someone last showed it working.

the other end

When the Allowed Loss Was Zero

At a smart-parking startup I once moved the time-series database that recorded every parking event. It held every event the company billed from, so a lost record was money nobody could bill.

The move was the riskiest step of the whole project, so it got the most rehearsal: a staged cutover, integrity checks, and performance baselines taken before and after. It landed with no data lost and performance where it had been.

Put in the two numbers, the allowed loss was zero, and those three habits are how you find out that zero held before anyone depends on it. The line I use for it is short: every parking record is an invoice; you do not get to lose one in transit.

Read as recovery objectives, my two examples are opposite sentences. The Redis comment says everything may go, because the truth is kept elsewhere. The migration says nothing may go, because this store is where the truth is kept. Every store you run sits between them, and each needs someone to place it on that line and say why.

the demo

What My Demo Would Answer

So I asked the demo the questions a reviewer would ask of each store: what may be lost, how long may recovery take, and where is the proof. My blueprint has a row for exactly this in its readiness list: backups restored in a drill. A row is ticked only with a dated run or drill log, and the list asks for the recovery time after losing a zone and after losing a region to be stated separately, each with the drill that proved it.

I searched the repository for the words backup and restore. They appear only in the blueprint's own instructions, the documents that tell a project what it has to contain. The compose file, the configuration, the scripts and the tests never use them. I also looked for a test that stops Postgres or empties Redis and then counts what is left, and found none, although the blueprint asks for one: kill the cache in the middle of a test and prove that idempotency still holds through the database.

Five recovery answers in the demo, graded with the four verdicts from Not Run Is Not Passing
The questionWhat the demo hasVerdict
May Redis lose everything?A sentence in the compose file with the reason, and a comment in the code that backs itpartial
Does Postgres keep what it acknowledged if its process dies?The default commit setting, and no test that kills itpartial
Does the data survive losing the volume?One volume and no second copynot done
Has a restore ever been timed, with a dated log?No log, and no backup to restorenot done
Are an RPO and an RTO written down for the store that holds the money?Nonot done

Nothing earns a pass. The two partials are answers that exist on paper, a comment for Redis and a default for Postgres, and neither has been tried.

the proof

A Number Is Real When It Has a Date

AWS's paper says a backup strategy has to include testing the backups, and my blueprint says the same thing as a tick that needs a dated log. A drill has a short recipe. Restore the newest backup into an empty environment, start the application against it, and write down four things: the date, the size that came back, the minutes it took, and the age of the newest record.

The last two are the numbers. The minutes are a recovery time you can now state without guessing, and the age of the newest record is the recovery point you actually got, which can be older than the one anyone promised. Run it for each failure you named, one zone, one Region, one corrupted table, and you have the separate answers the blueprint asks for.

Until then, the honest verdict for the store that holds the money is not done, and the Redis comment shows how little it takes to reach partial: one sentence with a reason in it.

// the part worth keeping

A recovery objective is a sentence with a reason in it.

The demo has no customers and holds no real money, so nothing was at risk here except the checklist. I would rather find an empty row in my own project than in a review of someone else's. The comment about Redis took one line to write. The Postgres service needs a backup, a restore and a date, and I have not written any of the three.
// carry forward

This page covers how much a system may lose, and how to show it. How slow it may be, and for how long, is a separate promise with numbers of its own.

// continue exploring