Cluster Restore
By the end of this page you will have restored every node of a ScramDB cluster to one consistent point in the past, the latest point the archive holds or a moment you choose, and started the restored nodes as a new cluster.
A cluster restore brings every shard group, and the cluster's own metadata (its tables, placement and node numbers), back to ONE commit position shared by every group. A transaction that wrote several groups is restored on all of them or on none, so the restored cluster never holds half of a transaction, whichever point you pick. It is the cluster's counterpart of point-in-time recovery on a single node, and it reads the same archive destination.
What the archive holds
With [storage.wal.archive] enabled on every node, each shard group's leader ships the group's committed log to the archive destination, and the leader of the cluster's metadata group ships the metadata log. Only committed log entries are shipped, and each segment is stored under a key only one leader can create, so a node that lost its leadership can never overwrite or interleave what the new leader shipped. When a group's log was compacted past what the archive holds (archiving switched on for a running cluster, for example), the group's leader ships a base: the group's state and a full copy of its rows, and the log continues from there. Each node's own storage log goes to its own place in the same destination, never mixed with another node's.
Point every node at the SAME destination, typically an object store, so the archive holds every group whichever node led it:
[storage.wal.archive]
enabled = true
destination = "s3://my-bucket/scramdb-archive"
# retention = "168h" # keep what a restore to any point of the last seven days needs
A destination left unset archives into each node's own data directory, which holds only the groups that node led: a cluster restore needs one shared destination. The archive's own metrics (scramdb_cluster_archive_* on Observability) show what each node ships, and scramdb_cluster_archive_gaps_total above zero says a group's history has a hole a restore cannot cross.
Choosing a target
| Flag | Restores to |
|---|---|
| none | The latest point every group of the archive can be restored to. |
--target-time <TIMESTAMP> | The state as of that time: every transaction acknowledged by then, and none after (2026-09-24T10:00:00Z, or 2026-09-24 10:00:00 read as UTC). |
--target-position <POSITION> | Exactly that commit position. |
--target-name <NAME> | A named restore point, made on the cluster with scram.create_restore_point (below). |
Give at most one. A time is turned into a point through a time index the cluster records as it commits: every node notes the commits it acknowledges, and every shard group's leader notes how far it has applied its log. The clock only helps you name the moment; it never decides what the restored point contains, which comes from the groups' own logs: a shard group that commits less often than the others gives the restore nothing it committed after the time, and a transaction that wrote several groups is restored on all of them or on none. Transactions and metadata changes (a CREATE TABLE, an ALTER TABLE) are placed against the target to within one millisecond.
Naming a restore point
Before a risky change, name the moment on any node, as a superuser:
SELECT scram.create_restore_point('before-migration');
Expect: one row with the commit position of the point. Every shard group is fenced at that position as the point is made, so a restore to it (--target-name before-migration) holds exactly what committed before it. A name is at most 63 bytes. Making a point under a name already taken keeps the first point, and answers its position; a restore to the name reaches that first point. PostgreSQL's own SELECT pg_create_restore_point('before-migration') makes the same point and answers its position in pg_lsn form.
Step by step
-
Plan the restore, once. From any machine that can read the archive:
scramdb restore-cluster plan --archive s3://my-bucket/scramdb-archive \--target-time "2026-09-24T10:00:00Z"Expect:
restore <ID> of <STREAM>: every group at commit position <P>, metadata through log index <M>, followed by the exact command to run on every node. For--target-timethe plan says what it holds on every group:restore <ID> of <STREAM>: the target time's cut at commit position <P>, <N> transaction(s) in flight at the time admitted, metadata through log index <M>, then one line per shard group,group <G>: log index <I> (base <B>), the log position the group had applied by the time (ornothing applied by the time). The plan is recorded in the archive, so every node restores to the same point even while the archive keeps growing.--streamnames the archived cluster when the archive holds more than one (the command lists them). -
Restore every node, before it starts. On each node, with the node's own configuration file and an empty data directory:
scramdb restore-cluster node --config /etc/scramdb/config.toml \--archive s3://my-bucket/scramdb-archive --stream <STREAM> --restore <ID>Expect:
node <NAME> restored to commit position <P>: <N> group(s), <R> row(s). The node gets the restored tables first, then the rows of every shard group it holds a replica of (a learner's into its columnar store), each group's log, and the cluster metadata. A data directory that is not empty is refused: restore into a fresh one. -
Stop every node of the old cluster (or keep it on a network the restored nodes cannot reach), then start the restored nodes as you normally start them. They form a new cluster at the restored point: every group's log starts with a marker at that position, so every new commit lands after every restored row, and the restored cluster archives into a new stream that records the one it came from and the point it was restored to.
Verifying you landed where you meant
Query for state that differs on either side of your target, on more than one node: a row committed just before it that must be present, a change just after it that must be absent, and a total that must match what you knew at that moment. Because every group is restored to one position, a transaction you know committed before the target is present on every node whole, and one that committed after it is absent everywhere.
Next
- Point-in-time recovery: the single-node form of the same idea.
- Observability: the cluster archive's metrics.