Skip to main content

Storage Tiering

By the end of this page you will know which storage-tier levers in ScramDB move data or bytes, how the cold tier moves sealed table data to object storage, and what it does not do yet.

Hot: the buffer pool​

The buffer pool is the real, working "how much of the database stays in fast memory" lever. It is covered in full on the Memory page, this section is a pointer, not a duplicate.

[storage]
buffer_pool_size_bytes = "8GB"
buffer_pool_percent = 60
buffer_pool_cap = "64GB"

Data inside the buffer pool is served from RAM. Data outside it falls back to a disk read. This is the tier that actually determines whether your working set feels fast, and it is the first thing to size correctly, see Memory for the full field reference and a verification recipe.

Warm: the OS page cache​

The [storage.memory] effective_cache_bytes / effective_cache_percent fields tell the planner how much data the OS page cache and the buffer pool are likely to keep warm, as PostgreSQL's effective_cache_size does. The index scan cost reads it: a row fetched again while the cache still holds it is not charged a second random read, so a larger value makes an index scan that returns many rows cheaper against a sequential scan.

The real "warm" tier is whatever your operating system's page cache keeps resident from recent disk reads, on its own, using its own eviction policy. ScramDB does not manage or control this layer directly, it is a side effect of the OS caching pages it has recently read from disk, and it disappears whenever the OS decides to evict a page for something else.

Off-box durability at object-storage cost: WAL archival​

This is a genuine, working cost lever, but it is a durability and point-in-time-recovery feature, not a way to move table data to cheaper storage. Do not confuse it with "cold" table tiering below, they solve different problems.

[storage.wal]
[storage.wal.archive]
enabled = true
destination = "s3://my-bucket/scramdb-wal"
poll_interval = "5s"
KeyDefaultWhat it controls
enabledfalse in the library default config; true in the shipped Docker image's default config. These differ, check which one applies to your deployment.Turns WAL segment archival on.
destinationunsetAn object_store URL: file:///abs/path, s3://bucket/prefix, gs://bucket/prefix, or az://account/container/prefix. If enabled = true with no destination set, ScramDB derives a local file archive under {data_dir}/wal-archive so point-in-time recovery and branching work out of the box with no object store configured.
poll_interval"5s"How often the archiver checks for newly closed WAL segments once it is caught up. This does not throttle a backlog, consecutive already-closed segments ship back to back with no delay between them.

Once enabled, ScramDB ships each closed WAL segment to the configured destination. This is what backs point-in-time recovery and database branching without requiring you to keep every WAL segment on local disk indefinitely, off-box object storage is cheaper than growing local disk to hold the same history. See Backup for how archived WAL is used for recovery and branching.

Verifying it worked: check the WAL-archive counters on /metrics.

curl -s http://127.0.0.1:9090/metrics | grep scramdb_wal_archive

Expected: scramdb_wal_archive_enabled 1, and scramdb_wal_archive_segments_archived_total climbing as writes generate and close WAL segments. scramdb_wal_archive_segments_behind and scramdb_wal_archive_bytes_behind only appear once a real measurement exists, their absence on a fresh scrape means "not measured yet," not zero.

Cold: table data in object storage​

The cold tier copies a table's sealed segments (its data files that no longer take writes) to object storage, removes their local files when the server next starts, and reads each one back to local disk the first time a query reads it. Local disk then holds the data in use, and the object store holds the rest.

[storage.cold_config]
cold_enable = true
destination = "s3://my-bucket/scramdb-cold/node-a"
cold_after_secs = 86400 # copy a sealed segment once it is untouched for a day
writeback_frequency_secs = 300 # look for segments to copy every five minutes
KeyDefaultWhat it controls
cold_enablefalseRuns the cold tier.
destinationunsets3://bucket/prefix, gs://bucket/prefix, az://container/prefix, file:///path or a plain directory. Unset, it is derived from cloud_api, bucket_name and base_path. Give every node a prefix of its own.
cold_after_secs86400How long a sealed segment stays untouched before it is copied.
writeback_frequency_secs300How often the tier looks for segments to copy.
concurrent_requests100The most requests in flight against the destination at once.

How it behaves:

  • Copies. A background thread, at idle disk priority, copies each sealed segment untouched for cold_after_secs, streaming it with a checksum, and writes a small marker file beside it naming the copy. A segment written to after its copy is copied again.
  • Local files leave at a restart. When the server starts, after it replays its log and before it opens any table, it removes the local file of every copied segment that has not changed since its copy. A running server does not remove files; the disk a copy frees comes back at the next restart.
  • Read back on first use. The first read of a segment whose file is away downloads it into place, checked against the checksum in its marker before it is used, so a query answers exactly as it did before. The segment then stays local until it is untouched for cold_after_secs again.
  • Bounded and counted. Every copy and read-back has a deadline that grows with its size and at most three attempts. A read-back that cannot get the right bytes fails the query that needed it, with the reason; nothing is ever read from a copy that does not match.
  • Nothing is deleted from the store. A backup of the data directory carries the marker files, and a server restored from it reads the segments back from the same store (configure the same [storage.cold_config]; a server without it says so when a query reaches a cold segment).

What it does not do yet: remove local files while the server runs, keep a bounded local cache of read-back segments, delete copies nothing refers to any more, or read part of a segment instead of the whole file.

Verifying it worked: the cold tier's counters on /metrics.

curl -s http://127.0.0.1:9090/metrics | grep scramdb_cold_

Expected: scramdb_cold_segments_demoted_total climbing once segments are older than cold_after_secs, scramdb_cold_segments_evicted_total after a restart, and scramdb_cold_segments_read_back_total as queries reach cold segments. The failure counters (scramdb_cold_demotion_failures_total, scramdb_cold_read_back_failures_total, scramdb_cold_read_back_mismatches_total, scramdb_cold_transfer_timeouts_total) stay at 0 on a healthy tier.

Next​

See Memory for the full buffer-pool field reference, or Backup for how WAL archival backs recovery and branching.