Skip to main content

Backups

By the end of this page you will have taken a base backup of a ScramDB instance, verified it landed correctly, and (optionally) configured your instance so that backup location also powers branch-from-timestamp.

What a base backup is​

A base backup is a cold, consistent snapshot of a stopped instance's data directory, plus a manifest recording exactly where in the write-ahead log (WAL) the snapshot was taken. scramdb backup opens the instance the same way the server itself starts up (full recovery and reconciliation, no shortcuts), triggers a checkpoint to get a clean, well-defined WAL position, then streams every file in the data directory except the live wal/ subtree to your chosen destination. The manifest is written last, so a backup interrupted partway through leaves either a complete backup or an incomplete one with no manifest, never a false-success partial backup that restore would accept.

This is a cold backup​

scramdb backup is explicitly offline. It refuses to run if the target instance's data directory is currently locked by a running scramdb server, with an error that says exactly that:

cannot back up '<dir>': the storage directory is locked by a running scramdb server.
This is a cold backup tool ... stop the server first, or back up from a stopped replica.

Online (hot) backup against a live, traffic-serving server is a separate capability, not built into this tool today. In practice this means one of two things:

  • Stop the server for the duration of the backup, then start it again, or
  • Point scramdb backup at a stopped replica or a copy of the data directory that is not currently held by a running instance.

Take a backup​

  1. Stop the server (or make sure you are pointing at a stopped copy of the data). If you are running the Docker image:

    docker stop scramdb
  2. Run scramdb backup, pointing --config at the config file whose storage.basedir names the instance you want to back up, and --out at a destination:

    scramdb backup --config /etc/scramdb/config.toml --out file:///var/lib/scramdb/backups/2026-08-02

    --out accepts a local directory path or an object storage URL: s3://bucket/prefix, gs://bucket/prefix, az://account/container/prefix, or file:///abs/path. A bare string with no scheme is treated as a local filesystem path.

  3. Check the output. A successful backup logs a single summary line:

    backup complete: N file(s), base_lsn=<LSN>, base_wal_file_id=<ID>, destination='file:///var/lib/scramdb/backups/2026-08-02'

    base_lsn is the exact WAL position the backup represents. Note it, or the timestamp you ran the backup at, you will want one of them if you ever restore or branch from this backup with an explicit target.

  4. Inspect what landed at the destination:

    ls /var/lib/scramdb/backups/2026-08-02

    Expect a manifest.json file and a data/ subdirectory holding the snapshotted tree. If manifest.json is missing, the backup did not complete; do not trust the data/ contents.

  5. Restart the server if you stopped it for the backup:

    docker start scramdb

If it fails: the most common cause is the storage-directory-locked error above, which means the instance is still running against that basedir. Stop it (or all containers/processes sharing that volume) and retry. A permissions or connectivity error against an object storage destination surfaces as a plain I/O error naming the destination; check your credentials and bucket/container name.

What's included, and what isn't​

Included: every file under the data directory except the live WAL subtree, this covers the catalog and all table data as of the checkpoint the backup took.

Not included: the live WAL segments. Those are handled separately, either from the local wal/ directory (bounded by your retention settings) or from the WAL archive, at restore time via --wal-source. See Restore for how the two come together.

The manifest also records a format version, the timeline the server was writing when the backup was taken, and that timeline's history: the timelines it descends from and the WAL position where each one began. A restore starts a new timeline, so the WAL of a restored server never collides with the WAL of the server it came from; see Timelines for how restores, archives and branches follow them.

A backed-up server keeps what its restore needs​

scramdb backup marks the server it backs up with a base_backup.v1 file in its wal/ directory, written before the manifest. From the server's next start on, it logs what a replay from that backup needs, the same way it does with a backup location or a WAL archiver configured: every compaction (the background one and VACUUM FULL), every ALTER TABLE that rewrites a table's rows and every COPY writes its page images into the WAL, so a restore or a branch to a moment after any of them holds exactly the rows the server held then. Expect the WAL, and the archive if you run one, to grow by about the size of the data those operations write.

A server restored from a backup is marked the same way, because a later restore of it replays its WAL after its own history. Leave the file in place: without it, and with no backup location and no archiver configured, those operations skip their page images again, and a restore or branch to a moment after one of them cannot rebuild those pages.

Configure a backup location for branch-from-timestamp​

CREATE DATABASE ... FROM ... AT TIMESTAMP ... (see Branching) lets a running server fork a database as of a past moment, without you invoking scramdb backup and scramdb restore yourself. To make that work, tell the running server where to find its own base backup:

[storage.wal.backup]
destination = "file:///var/lib/scramdb/base-backup"

A branch replays the server's own WAL from that backup, so with a location set every COPY keeps its page images in the local WAL, even with [storage.wal.archive] disabled, and a branch at or after a COPY's commit holds its rows. The WAL grows by about the loaded data's size for it. A server with no backup location and no archiver skips those images until it has been backed up (see A backed-up server keeps what its restore needs).

This is one configured location holding the current base backup, not a searchable registry of several dated backups. If you take a fresh backup to this destination, branch-from-timestamp and PITR-via-archive get that backup's coverage; picking automatically among multiple historical backups is not something ScramDB does today, so refresh the backup at this location on whatever cadence your recovery window requires.

To actually populate that location, run scramdb backup with --out pointed at the same destination:

scramdb backup --config /etc/scramdb/config.toml --out file:///var/lib/scramdb/base-backup

WAL archiving is the backup's companion​

A base backup only gets you back to the moment the backup was taken. To restore or branch to a moment after that, ScramDB needs the WAL written since then. Without WAL archiving, only the WAL still sitting in the instance's local wal/ directory is available, and that is bounded by your retention settings, older segments get pruned. With archiving enabled, WAL is shipped continuously to a durable location that restore can pull from, extending how far forward you can replay from a given backup.

The shipped Docker image enables archiving by default. See Point-in-time recovery for the full [storage.wal.archive] and [storage.wal.retention] reference.

Next​

  • Restore: rebuild a fresh instance from this backup.
  • Point-in-time recovery: restore to an exact moment instead of the latest available point.
  • Branching: use this backup to fork a database as of a past moment, on a live server.