Run restore drills and read the recovery evidence
Prove that a stored backup can be restored in isolation, and read the reconciliation of rows, file digests and ledger totals.
Before you begin
A backup that has never been restored is an unproven backup. A restore drill restores a stored backup into an isolated scratch copy, compares it with what was backed up, and removes the copy. It also feeds the activation gate of releases.
You need a successful, stored backup. Take and verify one under Administration > Settings > Storage and backup. Backups taken before file and ledger fingerprints were recorded can only be checked by row counts.
Run a drill
- Open Platform > Operations > Restore drills.
- Press Run a restore drill.
- Choose the Backup: Newest stored backup, or one of the last 50 successful backups. If none exists: 'Choose a stored, successful backup.'
- Press OK. The drill (
DRL-00001) is queued. - Open Platform > Operations > Job runs and press Run the scheduler now.
- Open the drill and press Refresh.
Read the result
A passed drill reads, for example, 'Restored n rows in m tables; every count, file digest and ledger total matches.' The reconciliation shows:
- Rows restored against rows expected, with a table of differences where they differ;
- stored files that match their digests;
- ledger totals that match;
- the RTO measured (the time the restore took) and the RPO measured (the age of the backup in minutes);
- the schema before and after.
A failed drill says why. Examples:
| Message | Meaning |
|---|---|
| 'The stored backup is not the file that was written: its checksum differs.' | The backup file was changed or damaged. |
| 'The backup was 1500 min old, beyond the RPO of 1440 min.' | The backup is older than the environment allows. |
| 'The restore took <n>s, beyond the RTO of 240 min.' | The restore was too slow. |
| '... This backup was taken before file and ledger fingerprints were recorded, so only row counts were reconciled.' | An old backup. Files and ledger show 'not fingerprinted'. |
A drill started from the Restore drills screen is not linked to an environment, so it is not measured against RTO and RPO targets, and it never fails on them. Drills started by an upgrade rehearsal from a release can be linked.
On PostgreSQL the database role must be allowed to create a scratch database. If not, the drill fails with 'Could not create the scratch database for the drill (...). The database role needs CREATEDB.'
Read the recovery evidence report
Open Platform > Reporting > Recovery evidence. It lists every drill with its scope, result, snapshot, time, rows restored against expected, file digests, ledger totals, RTO and RPO in minutes. The line at the top reads, for example, 'n / m drills passed - x with every row count matching - y with file digests matching - z with ledger totals matching.' The word Reconciles shows green only when every finished drill passed. Use the list tools to export the same rows to a file.
Upgrade rehearsals
A rehearsal is a drill that also runs the release's migrations on the restored copy. It is started from a release (see Prepare and submit a release) and shows in the same list. A passed rehearsal must have Reached the release's migration head.
Good to know
- A drill job cannot be scheduled in the job engine yet. Run drills by hand.
- The alert policy DRILLS raises a warning when any drill failed in the last 30 days, and BACKUP-AGE raises a critical alert when the last good backup is older than 26 hours. See Alerts and workload.
- The Overview tile Last good backup shows the age in hours, red when there is none or it is older than 26 hours.