Skip to main content

Config file

By the end of this page you will have the complete list of every section and key ScramDB's TOML config file accepts, each with its type, real default, and effect, so you can look up any setting without guessing.

File format​

The config file is TOML. Every setting lives under a [section] table (or a nested [section.subsection] table); there is no flat top-level key. An unrecognized top-level key is a hard parse error naming the exact field, so a typo fails loud at startup rather than being silently ignored.

Two value grammars are reused across almost every section:

  • Byte sizes accept a plain integer (bytes) or a human-readable string: "4GB", "512MB", "768mb" (the suffix is case-insensitive), "1TB", "16KB". Recognized suffixes are B (or no suffix), KB/K, MB/M, GB/G, TB/T, PB/P, all 1024-based. An unrecognized suffix is a hard parse error naming it, and so is a negative size or one too large to count in bytes.
  • Durations accept a human-readable string ("30s", "5m", "200ms", "24h") or a plain integer of milliseconds.

A few fields use the string "auto" as a sentinel meaning "detect and size this automatically"; those are called out individually below.

Every section is optional. Omitting a section is byte-for-byte identical to writing that section out with every field at its default. The sections below are ordered the way a real config file reads top to bottom: general, storage and its sub-tables, JIT, GPU, UDF and its sub-tables, cluster and its sub-tables, then the query-execution sections.

[general]​

The one section that holds server-wide knobs with no home of their own; everything else lives under its own named section.

KeyTypeDefaultEffect
jit_enabledbooltrueEnable LLVM JIT compilation for hot bytecode regions.
metrics_portu169090Prometheus /metrics and plain /health HTTP port. 0 disables the server entirely. A CLI flag can override this; see Precedence.
pg_addressstring, host:portunset in the engine (falls back to the CLI flag's own default, 127.0.0.1:5432, loopback-only)The PostgreSQL wire-protocol listen address.
pg_no_authboolfalseDisables PostgreSQL client authentication entirely and ignores hba_file while it's on.
shutdown_drain_timeout_secsu64 (seconds)10Budget for active queries to finish during a graceful shutdown (SIGINT/SIGTERM).
shutdown_jit_quiesce_timeout_secsu64 (seconds)10Bound for waiting out an in-flight background JIT compile at shutdown.
hba_filestring (path)unsetPath to a pg_hba.conf-format access-control file. Unset uses the built-in default rules (localhost trusted, password elsewhere). Ignored entirely when pg_no_auth is set.
tls_certstring (path, PEM)unsetTLS certificate for the pgwire listener. Unset (either alone) means TLS is disabled (plaintext).
tls_keystring (path, PEM)unsetTLS private key. Must be set together with tls_cert to enable TLS.
license_keystringunset (Community edition)Enterprise or Trial license token. SCRAMDB_LICENSE_KEY takes precedence when both are set; see Environment Variables.
max_connectionsinteger, ≥ 1fitted to the machineClient connections served at once (PostgreSQL's max_connections). Each connection place reserves a little memory for the life of the process, so the places are part of the memory budget the other budgets are fitted around. Absent, it is two percent of the memory envelope divided by what one place reserves: never below PostgreSQL's 100 on a machine with room for them, never above what the process's open-file, thread and memory-map limits allow (a machine too small for 100 gets what fits, with a warning). A value the kernel's limits or the memory envelope cannot hold is refused at startup, naming the arithmetic; it is never shrunk.
superuser_reserved_connectionsinteger, below max_connections3Connection places only a superuser may take (PostgreSQL's superuser_reserved_connections).
authentication_timeout_secsu64 (seconds), > 060How long a client may take from connecting to finishing authentication (PostgreSQL's authentication_timeout) before its connection is closed.

The 0.0.0.0:5432 binding you get from docker run scramdb/scramdb:latest comes from the Docker image's own startup command, not from ScramDB's own default, which is loopback-only. If you're running the binary directly and want it reachable from outside the host, set pg_address explicitly.

An unreadable hba_file, or a tls_cert/tls_key pair that can't be turned into a working TLS listener (unreadable file, bad format), is a fatal startup error, never a silent fallback to defaults or to plaintext.

[storage]​

Three fields have no default and must be set: prod_name (string), shard_id (a non-negative integer), and basedir (path). A config file that omits any of the three fails to parse.

KeyTypeDefaultEffect
data_dirpath{basedir}/data (derived if empty)Segment file directory.
buffer_pool_size_bytesbytes or string0 (unset sentinel, auto-sized from buffer_pool_percent of detected RAM, capped at buffer_pool_cap)Buffer pool size. A non-zero value pins an exact size and skips auto-sizing.
buffer_pool_percentu8, 0-1000Percent of detected RAM the auto-sizer targets when buffer_pool_size_bytes is 0. 0 auto-derives the percentage from the active I/O backend; set 1-100 to target that percent of detected RAM yourself.
buffer_pool_capbytes/string"64GB"Ceiling the auto-sizer never exceeds regardless of detected RAM.
memory_hard_limit_fractionf64, [0.1, 0.95]0.8Hard ceiling on process RSS as a fraction of detected RAM; a breach cancels active queries loudly.
execution_memory_bytesbytes/string0 (unset sentinel, auto-sized)Per-query execution budget (hash tables, sorts, morsel buffers); a query over budget spills to temp_dir. A non-zero value pins an exact size and skips auto-sizing.
execution_memory_percentu8, 0-1000 (off; execution_memory_bytes governs)When above 0, wins over execution_memory_bytes, auto-sized from this percent of detected RAM, capped at execution_memory_cap.
execution_memory_capbytes/string"64GB"Ceiling for the execution-memory auto-sizer; consulted whenever the budget is auto-sized, meaning execution_memory_percent is above 0, or neither execution_memory_percent nor execution_memory_bytes is set.
temp_dirpath{basedir}/spill (derived if empty)Spill-file directory.
segment_max_rowsu641000000Rows-per-segment rotation threshold.
flush_threshold_rowsusize8192Rows buffered per table before an active segment flushes. Must be at most segment_max_rows.
prefetch_depthusize8Pages the sequential-scan prefetch worker reads ahead. Must be at most prefetch_queue_depth.
prefetch_queue_depthusize64In-flight prefetch request queue depth.
bloom_max_bytesbytes/string"1KB"Max serialized bytes for one column's page-level Bloom filter.
wal_dirpath{data_dir}/wal (derived if empty)Write-ahead log directory.
copy_staging_limitbytes/string, ≥ 1, optionalhalf the data volume's free space at start (half the execution memory budget where the volume cannot be measured)Most bytes of COPY rows staged outside their tables before their transactions commit. A COPY that would stage past it fails with a disk-full error (53100) and stages nothing more.

[storage] also validates several cross-field constraints at config load and fails loud, naming the offending values, if any are violated: buffer_pool_percent and execution_memory_percent must be in range, flush_threshold_rows must not exceed segment_max_rows, prefetch_depth must not exceed prefetch_queue_depth, and if buffer_pool_size_bytes is set explicitly, buffer_pool_size_bytes + execution_memory_bytes must not exceed detected total RAM.

[storage.memory]​

Postgres-class per-component memory budgets. All sizes are byte strings.

KeyTypeDefaultEffectStatus
per_operator_bytesbytes/stringunsetPostgreSQL's work_mem: the memory each sort, hash join build, hash aggregation and window input sort holds before it spills to disk. Unset, each operator holds its share of its statement's execution memory instead. 64kB to 2147483647kB. A session sets its own with SET work_mem.Wired.
maintenance_bytesbytes/string0 (auto-sized to a third of the resolved execution memory pool)ANALYZE / CREATE INDEX working-memory ceiling, reserved whole from the shared execution memory pool. A value that cannot fit the execution pool it's reserved from is capped to what fits, with a warning, rather than refused.Wired and enforced today.
hash_mem_multiplierf64, 1 to 10002.0PostgreSQL's hash_mem_multiplier: a hash join build or hash aggregation holds per_operator_bytes times this before it spills. Applies only when per_operator_bytes (or a session's work_mem) is set. SET hash_mem_multiplier per session.Wired.
working_set_bytesbytes/string0 (auto-sized from working_set_percent)The envelope every memory budget in this table and in [storage] is fitted into when the server starts; a combination of explicit sizes that cannot fit refuses to start, with the arithmetic.Wired.
working_set_percentu8, 1-10075Percent of RAM targeted when working_set_bytes is 0.
read_buffer_pool_bytesbytes/string0 (auto-sized from detected RAM)Hard cap on the page pool's total working set (checked out plus idle-recycled).Wired and enforced today: the pool backpressures rather than growing past it, and is raised automatically with a loud log line if configured below the pool's own deadlock-free floor.
effective_cache_bytesbytes/string0 (auto-sized from effective_cache_percent)How much of the data the planner should assume cached (PostgreSQL's effective_cache_size).Read by the index scan cost.
effective_cache_percentu8, 1-10050Percent of RAM targeted when effective_cache_bytes is 0.
heavy_concurrencyusize, ≥ 11How many heavy statements the admission ledger admits at once. Each admitted statement gets execution_memory_bytes / heavy_concurrency. A nested statement (a scalar subquery, an uncorrelated IN (subquery)) inherits its ancestor's grant instead of requesting a second one, so heavy_concurrency = 1 does not deadlock against a statement's own nested statements.Wired and enforced today.
admission_wait_warn_msu64, ≥ 130000How often a heavy statement waiting for admission logs a warning. This is a logging cadence, never a refusal deadline: a parked statement is never timed out.Wired and enforced today.

Set a *_bytes field to pin an exact value, or leave it at 0 to auto-size from the matching *_percent of detected RAM. effective_cache_bytes is read by the index scan cost, as PostgreSQL reads effective_cache_size.

[storage.io]​

KeyTypeDefaultEffect
backend"auto" | "file" | "buffered" | "io_uring" | "nvme_passthrough""auto""auto" and "file" both resolve to the always-safe direct-I/O backend (O_DIRECT reads and writes) today; "auto" doesn't yet pick io_uring automatically. "buffered" is the buffered (non-direct) file backend, reading and writing through the OS page cache instead of O_DIRECT. "io_uring" is real and safe under concurrent workers. "nvme_passthrough" requires nvme_device and is validated present and usable at config-load time.
nvme_devicepath, optionalunsetThe exact NVMe character device (/dev/ngXnY) nvme_passthrough targets. Never auto-discovered: required, and probed, only when backend = "nvme_passthrough". A missing, wrong-type, or unusable device is a hard validation error naming the exact reason.

[storage.cold_config]​

The cold tier: a table's sealed segments copied to object storage, their local files removed at the next start, and each read back the first time anything reads it. Every key has its own default, so the table may set only the keys it changes. See Storage Tiering for how it works and what it does not do yet.

KeyTypeDefaultEffect
cold_enableboolfalseRuns the cold tier. A running tier with no destination, or with any of the three numbers below at 0, refuses to start with the reason.
destinationstring (URL or path), optionalunsetWhere copies go: s3://bucket/prefix, gs://bucket/prefix, az://container/prefix, file:///path or a plain path. Unset, it is derived from cloud_api, bucket_name and base_path. Give every node a prefix of its own.
cold_after_secsu64, ≥ 186400How long a sealed segment stays untouched before it is copied. A segment read back starts its age again.
writeback_frequency_secsu64, ≥ 1300How often the tier looks for segments to copy.
concurrent_requestsusize, ≥ 1100The most requests in flight against the destination at once.
cloud_api"GCP" | "AWS" | "Azure""GCP"The store bucket_name lives in, when destination is unset.
bucket_namestring""The bucket (or Azure container), when destination is unset.
base_pathstring"snapshots"The path inside bucket_name, when destination is unset.

[storage.wal]​

KeyTypeDefaultEffect
max_file_sizebytes/string"64MB"WAL file rotation size.
sync_every_writeboolfalsefsync after every append (per-commit durability, at a large latency cost) instead of relying on group-commit via epoch_duration.
epoch_durationduration"10ms"How often the flush worker batches and writes buffered WAL entries, even absent a size trigger. Must be greater than zero; a zero value would busy-spin the flush worker at 100% CPU.
flush_threshold_bytesbytes/string0 (automatic)Bytes of log records waiting in the WAL ring that trigger an early flush ahead of the epoch timer. The ring holds 64 of them, reserved from the memory budget. 0 sizes it for the machine when the server starts: a megabyte of ring per query worker, at most 1/256 of the working set, between 1 MiB and 256 MiB. A value pins it.
synchronous_commitbooltrueWhether a transaction commit waits for its WAL entries to be durable before acknowledging the client. false acknowledges without waiting.

[storage.wal.archive]​

KeyTypeEngine defaultEffect
enabledboolfalseStarts the WAL archiver task, which ships closed segments to a destination for point-in-time recovery and branching. false costs nothing: no task is spawned.
destinationstring (URL or path), optionalunsetAccepts file:///abs/path, s3://bucket/prefix, gs://bucket/prefix, or az://account/container/prefix. If enabled = true and this is unset, it derives to a local {data_dir}/wal-archive, so PITR and branching work out of the box on a single node.
poll_intervalduration"5s"How often the archiver checks for newly closed segments once caught up. A backlog ships back-to-back with no added delay regardless of this setting.
retentionduration"168h"The archive deletes segments by default: an archived segment is removed once it is older than this AND lies before the base backup's WAL position at [storage.wal.backup] destination; with no backup there, age alone decides. The newest segment always stays. Set a longer duration (hours, for example "720h" for 30 days) to keep more point-in-time history.

The engine's own default for enabled is false. Separately, every config file ScramDB ships as a shipped reference or example explicitly sets enabled = true, so anyone starting from one of those files gets WAL archiving on. Both facts are true at once; they aren't the same thing.

In a cluster, leaving destination unset is safe for a single node but risky once there's more than one: each node's local {data_dir}/wal-archive is node-local, so an unset destination fragments PITR history across leader failovers. Set destination to shared or object storage for any real cluster. See Clustering.

[storage.wal.checkpoint]​

KeyTypeDefaultEffect
enabledbooltrueRuns the periodic checkpoint scheduler: a checkpoint every interval, or sooner once the WAL has grown by wal_bytes since the last one. false leaves checkpoints to a clean shutdown and the CHECKPOINT command, and the WAL grows until then.
intervalduration"5m"Fire a checkpoint at least this often regardless of WAL volume (PostgreSQL's checkpoint_timeout).
wal_bytesbytes/stringadaptiveFire a checkpoint once the WAL has grown by this many bytes since the last one, even before interval elapses. Unset, a sixteenth of the free space on the WAL's volume, between 64MB and 1GB, taken again after every checkpoint, so a filling disk checkpoints sooner.

Periodic checkpoints are on by default, in the engine and in every shipped config: they keep a running server's WAL bounded and a restart's replay short. A checkpoint's disk writes run at idle priority until the WAL has grown past twice wal_bytes.

[storage.wal.retention]​

KeyTypeDefaultEffect
min_keptusize1Always keep at least this many closed WAL files, regardless of checkpoint coverage, archive state, or anchors.
archive_awarebooltrueOnce the archiver has started, a WAL file isn't droppable until the archiver's own watermark has passed it. false ignores archive state; only safe when point-in-time recovery via the archive isn't relied on.

[storage.wal.backup]​

KeyTypeDefaultEffect
destinationstring (URL or path), optionalunsetWhere a live instance finds its own base backup for from-timestamp branch materialization (CREATE DATABASE ... FROM ... AT TIMESTAMP and equivalent). Same accepted URL forms as [storage.wal.archive].destination. Unset means branch-from-timestamp refuses loudly, naming this exact key, rather than silently doing nothing. A set location also keeps the page images of every COPY in the local WAL, even with [storage.wal.archive] disabled, so a branch at or after a COPY's commit holds its rows; the WAL then grows by about the loaded data's size.

[storage.gc]​

KeyTypeDefaultEffect
cycle_intervalduration"30s"Interval between MVCC garbage collection cycles.
max_cycle_durationduration"5s"Max duration for a single, time-bounded, incremental GC cycle.
max_snapshot_ageu64 (timestamp units, not wall-clock)300Max snapshot age before a transaction is excluded from the GC watermark; older transactions get a snapshot-too-old error rather than blocking cleanup indefinitely.

[storage.compaction]​

KeyTypeDefaultEffect
target_segment_rowsu641000000Target rows per compacted segment.
min_segments_to_mergeusize4Minimum frozen segments needed to trigger a merge.
delete_fraction_thresholdf640.3Maximum fraction of deleted rows a segment may hold before it's forced into compaction regardless of the other thresholds.
check_intervalduration"60s"Interval between compaction checks.
maintenance_budget_bytesusize0 (derived)Memory the background compactor may use. 0 derives it as half the execution budget.

[jit]​

Compiled-artifact cache sizing.

KeyTypeDefaultEffect
compiled_cache_memorybytes/string"512MB"In-memory bound for compiled bytecode and JIT-tier kernels, and for the working memory of the compiles running now: cached code is evicted to make room for a compile.
compiled_cache_diskbytes/string"1GB"On-disk bound for the cache's compiled-artifact files.

[gpu]​

GPU acceleration. This section is only present on builds compiled with GPU support; check with your build or image before relying on it being available in your deployment.

KeyTypeDefaultEffect
enabledbooltrueEnables GPU acceleration. Safe to leave on with no GPU hardware present: ScramDB detects the absence, logs one informational line, and falls back to CPU-only execution rather than failing to start.
device_idi320GPU device index. -1 auto-selects the device with the highest compute capability.
vram_budgetbytes/string or 00 (use 90% of free VRAM)Maximum VRAM budget.
vram_hash_table_budgetbytes/string or 00 (60% of vram_budget)VRAM reserved for hash table builds.
pinned_memory_poolbytes/string"256MB"Pinned host-memory pool size, used only as an automatic fallback when on-demand pinning fails. No capability flag or privileged mode is required for ScramDB to run with a GPU by default; raising this value is only relevant in the narrow case where the automatic pinned-pool fallback itself also fails to allocate.
streams_per_workeru322GPU streams per worker thread, for double-buffering.
min_batch_sizeu64 (rows)50000Minimum batch size eligible for GPU dispatch.
kernel_timeout_msu64 (ms)5000Kernel timeout. 0 disables it.
vendor"auto" | "nvidia" | "amd" | "apple""auto"Preferred GPU vendor.
arch_overridestring"" (auto-detect)Architecture override.
metrics_enabledbooltrueIntended to publish GPU metrics on the Prometheus /metrics endpoint. No observable effect today: the counters exist but nothing reads them onto /metrics, so leaving this true adds no GPU series to the endpoint. Consistent with GPU tuning and Observability.
debug_launchesboolfalseLogs every GPU kernel launch at debug level.

Every shipped example or Docker config that touches [gpu] at all sets only enabled = true and leaves everything else at its default. See GPU Acceleration for the full picture of what runs on a GPU and when.

[udf]​

The afterburner UDF runtime.

KeyTypeDefaultEffect
enabledbooltruefalse makes CREATE FUNCTION / INSTALL MODULE for afterburner languages fail loudly; the shared runtime is never built, so there's zero footprint when off.
fuel_per_batchu64, > 0500000000Work (roughly, instructions) one batch of a function's calls may do before it is stopped: a coarse bound on a runaway function.
memory_bytesbytes/string, > 0"64MB"Per-invocation linear-memory cap; also the pooling allocator's per-instance hard cap.
timeout_msu64 (ms), > 01000Per-invocation wall-clock backstop. The underlying timer tick runs at a fixed 10ms granularity regardless of this value, so very small values won't get finer-grained enforcement than that.
output_bytesbytes/string, > 0"16MB"Per-invocation result-size ceiling.
engine_slack_bytesbytes/string"32MB"Sizing slack added over the measured engine-fixed memory footprint.
aot_cache_dirpath, optionalunset ({basedir}/udf-cache)The on-disk cache of compiled function modules. A relative path is placed under basedir, an absolute one used as written.
aot_cache_max_bytesbytes/string, > 0"512MB"The most bytes the compile cache holds; the least recently used modules are removed first, at start and after each new module is saved.
preinstall_dirstring (path)"/opt/scramdb/dist"Directory of curated packages installed automatically at startup, non-blocking, after the PostgreSQL listener is already up. Installation is idempotent across restarts. An empty string disables it. A host where this directory doesn't exist skips it silently, with only a debug-level log line.

[udf.registry]​

Governed access for name-form package install and upgrade.

KeyTypeDefaultEffect
urlstring (URL), validated"https://registry.afterburner.sh"Registry base URL for name-form install, upgrade, and version-listing calls.
token_filepath, optionalunset (anonymous access)Path to a bearer token file for authenticated registry calls, read fresh at each call. Never put the token inline in the TOML file. The file must be mode 0600 (not group or world readable); config validation fails loud, naming the exact mode found, if it isn't.
allowarray of glob strings["*"]Coordinate glob allow-list, checked before any network call.
denyarray of glob strings[]Coordinate glob deny-list, checked before allow; a match refuses loudly.
offlineboolfalsetrue makes every name-form install/upgrade/version verb fail loudly, naming this key, instead of making a network call. The bytes-upload install form always works regardless, since it performs no network call.
auto_update"off" | "check" | "apply""off""off": no ambient background task. "check": periodically resolves available upgrades into a view, but never applies them. "apply": additionally applies upgrades on the same cadence, honoring any version pins.
update_intervalduration, > 0 when auto_update isn't "off""24h"Cadence for the auto_update background task. Inert while auto_update = "off".

[udf.thrust]​

The governed worker pool UDF invocations run on.

KeyTypeDefaultEffect
enabledbooltruefalse runs scram.invoke / scram.invoke_batch inline on the calling task. true (the default) queues them onto this governed pool instead.
workersusize0 (auto: clamp(available_parallelism / 8, 1, 4))Pool worker count. This bounds isolation, not query throughput; query throughput comes from ordinary morsel fan-out, not from this pool.
nicei8, -20..=1919OS nice value for pool workers.
numa"auto" | "off""auto""auto" derives CPU pinning from real hardware topology, minus the cores reserved for query execution. "off" disables NUMA-aware pinning entirely (nice-value governance only).
affinityarray of core ids, optionalunset (auto-derive)Explicit CPU affinity mask, overriding numa-driven auto-derivation. Must be non-empty if set.
queue_depthusize0 (afterburner's own default of 256)Bounded work-queue depth for the pool.

[udf.daemon]​

Autostart for long-running packages. After the preinstall sweep, an installed package whose manifest declares [metadata.daemon] autostart = true and whose manifold grants listen is started in-process on the engine's own reactor, supervised, and surfaced in scram_daemons. The shipped consumer is scramdb/semantics, the Semantic AI MCP server on port 9191, which is why the release images publish that port. A distribution with no daemon-declaring package makes this a no-op scan.

KeyTypeDefaultEffect
enabledbooltrueMaster switch for the whole seam. false starts no daemons at all and logs the reason; nothing else about package installation changes.
shardsusize1Shards per daemon. Only HTTP daemons expand past 1.
env_allowlistarray of stringsPGHOST, PGPORT, PGUSER, PGPASSWORD, PGDATABASE, MCP_HOST, MCP_PORT, MCP_AUTH, SEMANTICS_MAX_ROWS, SEMANTICS_POOL_SIZE, SEMANTICS_MAX_CRED_POOLSProcess environment variables forwarded into daemon invocations. This is a ceiling, not a grant: the package's own manifold allowlist still decides what the script may actually read, so a variable must pass both to reach it.

A daemon that fails to start is a loud ERROR and a failed row in scram_daemons, never a wedged boot. A pool whose shards all die is restarted with bounded exponential backoff, base 2 seconds and doubling, up to 5 restarts; after that it stays failed with the reason rather than disappearing.

[cluster]​

Presence of a [cluster] table, even an empty one, activates cluster mode; absence is byte-identical to running as a single node.

A [cluster] table on a Community-edition node is a startup failure: ScramDB refuses to start the process at all unless the resolved license is Enterprise or Trial. This happens before any port is bound; it's a process exit, not a SQL-level error. See Clustering for the exact gate and how to supply a license.

KeyTypeDefaultEffect
node_namestring, optional, ≤ 255 bytesunset (falls back to the OS hostname)This node's stable name.
cluster_listenhost:port"0.0.0.0:7190"Address this node's cluster transport listens on for control traffic: membership, consensus, the commit protocol, forwarded DDL and lock calls, shuffle fetches. Seeds and advertise_addr name this port.
cluster_interactive_listenhost:port, optionalcluster_listen's address and the next port (7191)Where this node listens for interactive traffic from other nodes: replies to forwarded point reads. Peers learn it from the control connection, so it needs no entry in anyone's seeds; port 0 binds one the system picks.
cluster_bulk_listenhost:port, optionalcluster_listen's address and the port after that (7192)Where this node listens for bulk traffic: query data exchange, broadcast join sides and snapshot transfer. Same rules as cluster_interactive_listen. The three listeners must be three different ports.
advertise_addrstring (host:port or DNS name), optionalunset (auto-detected from the first non-loopback interface plus cluster_listen's port)Address other nodes dial to reach this node. Always set this explicitly behind NAT or in a container network.
seedsarray of host:port strings[]Static seed peers. Empty means a cluster-of-one: a fully functional single-node cluster with zero distributed overhead.
bootstrap_expectu32, optionalunset (seeds.len() decides)Seedless formation: how many other voter-role nodes must connect before this node founds the cluster, so a Helm chart or other zero-seed deployment can bring up every replica waiting on the same count instead of a static peer list. 0 founds a cluster of one. Refused at boot if it disagrees with a non-empty seeds list, or if it's positive with no seeds and no address-capable discovery provider (dns with dns_name set, or swim) that could ever supply the peers.
regionstring, optionalunset (regionless)This node's home region. When set, a table homed in this region keeps all of its voters here, so its transactions commit at region-local quorum latency instead of crossing the WAN. Advertised to peers and visible in scram.nodes. An empty or whitespace-only value counts as unset, never a region literally named "".
zonestring, optionalunsetA finer failure domain inside region (rack or availability zone). Reported in scram.nodes alongside region; placement treats it as informational today rather than a hard constraint.
columnar_replicaboolfalsetrue makes this voter also mirror every committed entry of the groups it owns into a full local columnar store, so whole-statement analytics run locally with zero network hops, while it keeps its normal bucket ownership. Costs one full columnar copy on this node plus a second apply per committed entry. Mutually exclusive with learner = true; the combination is refused at boot.
learnerboolfalsetrue makes this a read-replica-only (columnar learner) node: it hosts learner groups only, owns no buckets, and is never promoted to a voting member.
discoveryarray of "static", "dns", "swim"["static", "dns", "swim"]Active peer-discovery providers. An unrecognized name is a hard validation error.
dns_namestring, optionalunsetHeadless-service DNS name for the "dns" discovery provider. Required (and validated) when discovery includes "dns".
dns_refreshduration, > 0"5s"How often the "dns" provider re-resolves dns_name.
swim_gossip_fanoutusize3Live members the SWIM protocol broadcasts a direct gossip message to every tick, on top of probe piggyback. 0 disables the extra fan-out.
replication_factoru32, ≥ 1 (and ≤ cluster size once seeds is non-empty)3Replication factor for shard groups.
truncate_node_timeoutduration, greater than zero"120s"How long a cluster-wide TRUNCATE waits for one node to take its exclusive lock on the table (every transaction still using the table there finishing first), and again for that node to empty its copy. A node that runs out of it fails the statement, naming the node; a node whose coordinator stops mid-statement gives its lock back on its own after twice this long.
vectortableThe vector settings that only mean something on a cluster (recall measurement of index parts, whole-cluster statistics views); see Vector index configuration.
tls_cert, tls_key, tls_capath, optional (all three together, or all omitted)unsetOptional mutual TLS between cluster nodes. Independent of pgwire's own [general] tls_cert / tls_key.
connect_backoff_minduration, > 0 (and ≤ connect_backoff_max)"200ms"Minimum reconnect backoff.
connect_backoff_maxduration"10s"Maximum reconnect backoff.
send_queue_bytesbytes/string, ≥ 1"64MB"Byte bound on the outbound send queue of each peer connection, per traffic lane (credit-based backpressure).
recv_queue_bytesbytes/string, ≥ 1"64MB"Sizes the traffic lanes that [cluster.transport] leaves on "auto": control gets half of it, interactive a quarter, bulk all of it (at least the largest broadcast join side, 128MB, plus one frame). A lane's size bounds, across all peers, what the node holds received and not yet consumed, and what it holds queued to send; twice the lanes' total counts toward the process memory budget.
group0_bootstrap_timeoutduration, > 0"60s"How long a node waits at bootstrap for every configured seed to connect before forming the metadata group. Never falls back to a partial voter set on expiry; startup fails loud instead.
default_bucketsu32, ≥ 18Shard buckets a new primary-key table is auto-assigned at CREATE TABLE.
point_forward_timeoutduration"500ms"Bound on one forwarded point read's round trip before falling back to the general path.
route_ceiling_bytesbytes/string"64MB"Above this estimated result size, a statement never rides one replica; it scatters instead.
fragment_any_replicabooltrueWhere the fragments of a distributed read run. true places them on any replica that holds their buckets, learners first (in this node's region: a learner, then a follower, then the group's leader), spread evenly with this node first on a tie; a replica that cannot serve the read refuses and another replica reads its buckets. false keeps every fragment on the buckets' owners. Never changes an answer.

Every cluster node listens on three ports: 7190 for control traffic (cluster_listen), 7191 for interactive traffic and 7192 for bulk traffic, one connection per traffic lane between each pair of nodes, so a large transfer never delays a heartbeat, a vote or a point read. Open all three between the nodes. The cluster transport and the pgwire client port are separate listeners entirely.

Cluster sub-tables​

More sub-tables exist under [cluster] for fine-grained tuning of the underlying protocols. Each follows the same "absent table = every field at its default, byte-identical" contract as every section above, and each field has its own default, so writing only some of them leaves the rest untouched.

None of them need to be set for a cluster to run correctly. They exist for operators tuning failure-detection sensitivity or protocol timing under specific network conditions, and getting failure-detector timing wrong in either direction, too twitchy or too slow, has real availability consequences. Every field is listed below with its real default so you can reason about the tradeoff instead of guessing.

[cluster.swim]​

Failure detection: how fast a dead peer is noticed, and how hard the cluster tries not to be wrong about it.

KeyTypeDefaultEffect
ping_timeoutduration"500ms"Bound on a direct Ping's Ack before escalating to indirect probing. Lower = faster detection, at the cost of more false Suspect escalations on a slow-but-alive peer.
indirect_timeoutduration"500ms"Bound on an indirect (PingReq-relayed) probe before declaring Suspect.
indirect_probe_countinteger3Random relays asked to indirectly probe a non-responding target. Higher = more confidence before Suspect, at the cost of more probe traffic per suspected failure.
suspicion_minduration"1s"Floor of the dynamic suspicion timeout, reached at maximal corroboration. Lower = faster Dead declaration once every peer agrees, at the cost of a shorter window to catch a false positive.
suspicion_maxduration"5s"Ceiling of the dynamic suspicion timeout, used when no peer has corroborated yet.
dead_tombstoneduration"30s"How long a Dead member is retained before being reaped. A rejoin inside this window resumes at a fresh incarnation instead of a brand new join.
max_piggyback_updatesinteger6Cap on membership updates piggybacked per outbound message. Higher = faster dissemination per message, at the cost of larger frames.
retransmit_multiplierinteger4Multiplier in the retransmit-limit formula lambda * ceil(log2(n + 1)). Higher = more redundant retransmission of each update, so more reliable dissemination for more gossip traffic.
gossip_fanoutinteger3No effect here. This node always takes its fan-out from swim_gossip_fanout in [cluster] instead. Kept as a real field so the table still round-trips, not a live knob.
max_healthinteger8Ceiling of the local-health counter (Lifeguard LHM). Higher = more headroom to distinguish "very healthy" from "somewhat healthy" before probe timeouts stop scaling.
max_health_multiplierfloat4.0Timeout multiplier applied at maximum local health. Higher = more tolerance for a node under heavy local load before it starts falsely suspecting peers.
protocol_periodduration"1s"How often the SWIM driver's tick fires (probe, suspicion and reap pacing), independent of the timeouts above. Lower = faster detection and dissemination, at the cost of more background probe traffic.

[cluster.consensus]​

One instance, shared by the metadata group and every per-shard replication group.

KeyTypeDefaultEffect
election_timeout_minduration"1s"Lower bound of the randomized election and pre-vote timeout. Lower = faster failover after a leader crash, at the cost of more spurious elections under network jitter. See the warning below before lowering this.
election_timeout_maxduration"2s"Upper bound of the same randomized timeout. The randomized spread between min and max is what avoids split votes.
heartbeat_intervalduration"50ms"The leader's steady-state heartbeat and AppendEntries broadcast period. Lower = fresher followers and faster leader-loss detection, at the cost of more replication traffic at idle.
max_entries_per_appendinteger64Cap on log entries sent in one AppendEntries. Higher = fewer round trips to catch a lagging follower up, at the cost of larger per-message payloads.
max_append_bytesinteger (bytes)1048576 (1 MiB)Byte cap on one AppendEntries payload, applied alongside the entry-count cap above (one entry always goes, whatever its size). Replication rides the control traffic lane, where one append queues ahead of every group's next heartbeat and vote, so it stays small. At most 8388608 (half the default [cluster.transport] max_frame_size); a larger value is refused at boot.
max_inflight_bytesinteger (bytes)4194304 (4 MiB)Ceiling of the in-flight window: the most log entry bytes a leader keeps sent but unacknowledged toward one follower, per replication group. Below it the window adapts to each follower: twice the bandwidth-delay product the leader measures toward it (see adaptive_inflight_window), so a distant follower keeps a full pipe and a nearby one holds little. A full window pauses sends to that follower until it acknowledges; one entry larger than the whole window still goes, alone. Must be at least 1.
max_inflight_msgsinteger256The most AppendEntries messages a leader keeps sent but unacknowledged toward one follower, per replication group; the count half of the same window. Must be at least 1.
adaptive_inflight_windowbooltrueAdapt each follower's in-flight window to the bandwidth-delay product measured toward it, under max_inflight_bytes. false holds every window at max_inflight_bytes. Changes only how much a leader sends ahead, never what is replicated or committed.
max_pending_readsinteger1000Cap on concurrently outstanding read-index requests. Bounds memory on a leader that is stuck unable to reach a quorum, rather than letting it grow unbounded.
group0_commit_max_batchinteger16Group commit of the metadata group: the most already-queued operations one wake-up serves after the first before they are made durable with one write. Smaller than the shard groups' cap because a metadata operation can be a whole DDL apply. 0 serves one operation per wake-up.
shard_commit_max_batchinteger64Group commit of a shard group: the most already-queued operations one wake-up serves after the first, made durable with one append and one sync for the whole batch. 0 serves one operation per wake-up.
coalesced_heartbeatsbooltrueOne heartbeat frame per peer node each interval, carrying every group this node leads with that peer, instead of one heartbeat per group; a follower being repaired still gets its own. false sends every group's heartbeats separately.
heartbeat_flush_delayduration, below heartbeat_interval"200us"How long a heartbeat frame to a peer waits to gather more groups once a group's commit index advanced, a read needs confirming or a small append waits to ride it. The frame leaves at once while the connection to that peer has nothing queued. "0" sends every such frame at once.
quiescencebooltrueAn idle group (nothing in flight, every member caught up, everything committed) stops its heartbeats and its members' election timers until something happens, with one small refresh per peer node every election_timeout_min; a member campaigns again when its leader's node is lost or misses two refreshes. Takes effect only with coalesced_heartbeats.
carry_small_appendsbooltrueA leader sends an append no larger than the frame that would carry it alone inside its next heartbeat frame to that peer, so a tiny entry costs its bytes and not a frame, and a follower's reply to an append rides back inside the heartbeat answer frame it sends that leader. false turns both off. Takes effect only with coalesced_heartbeats.
leader_sends_before_syncbooltrueA leader sends new log entries while its own write of them is still syncing, and counts itself toward a commit only once synced; false syncs first, then sends. Never changes what commits.
threadsintegera quarter of the cores, 2 to 8Threads that drive the consensus groups: every replication group and the metadata group of the node runs on them, never a thread per group. Absent, they are sized from the cores like the threads that serve connections. 1 to 256.
log_segment_sizebytes/string, optionala thousandth of the data volume, 64MB to 256MBThe size each segment file of the node's shared raft log grows to before the next one starts. Every replication group the node hosts writes to this one log, so one sync covers them all; space is reclaimed a whole segment at a time. Absent, it is sized from the data volume (64MB when the volume cannot be measured). 1MB to 4GB.
Do not lower the election window without reading this

The election defaults are not conservative by accident. An earlier 150ms default caused election churn whenever synchronous applies starved the IO runtime, which turns a busy cluster into a leaderless one exactly when it is under load. Raise these values for WAN distances; lower them only with measurements in hand.

[cluster.consensus] was named [cluster.raft] before this release. The old spelling stays accepted forever as an alias, so a config file written against any earlier release keeps working with zero edits. A node started with it says so once in its startup log, and a file that sets both [cluster.raft] and [cluster.consensus] is refused, naming both (see Renamed keys). The new name is deliberately algorithm-neutral: every knob in this table is a property of leader-based replication in general, not of one specific consensus protocol, so a future change to the algorithm underneath ScramDB never forces you to rewrite this file. That is a compatibility promise: [cluster.raft] keeps working in every future release, full stop.

[cluster.rebalance]​

How quickly shard-group membership converges after a topology change.

KeyTypeDefaultEffect
catch_up_graceduration"200ms"Grace period after a learner is added before the driver tries to promote it. Lower = faster promotion, at the risk of promoting a learner that has not genuinely caught up. This is a bounded-wait heuristic, not a proof of catch-up.
retry_backoff_minduration"20ms"Floor of the retry backoff after a transient membership-change failure.
retry_backoff_maxduration"2s"Ceiling of that retry backoff.
reconcile_intervalduration"200ms"How often the periodic reconciler re-derives and converges each owned and led group's membership. Lower = a designated learner is added sooner after a table is created, at the cost of more frequent passes (cheap, and a no-op once converged). Must be greater than zero; 0 is rejected loudly at boot rather than panicking later.

[cluster.transactions]​

How long a COMMIT on cluster tables waits for the commit protocol. A commit is answered when every shard group it wrote holds its vote durably; these bound how long the session waits for that answer before it asks, and then for the answer to its question. An absent table is every key at its default.

KeyTypeDefaultEffect
commit_answer_timeoutduration, greater than zero"10s"How long a COMMIT waits for the commit protocol's answer before it asks for the transaction's outcome. Also how long a statement, right after its node restarted or right after a table it names was created, waits for the node to learn where that table is placed; past it the statement fails with SQLSTATE 57P03 having done nothing, and a COMMIT PREPARED leaves its transaction prepared (Two-Phase Commit).
outcome_query_timeoutduration, greater than zero"1s"How long that question waits for its answer. An outcome still unknown then is reported in doubt (SQLSTATE 08007): the transaction may still commit, so a client checks before it retries (see Error codes).

[cluster.transport]​

Connection establishment between nodes, and how fast a dead or stalled peer connection is noticed.

KeyTypeDefaultEffect
handshake_timeoutduration"10s"Bound on the Hello and HelloAck round trip, on the whole TLS-accept-plus-handshake span, and on the TCP connect and TLS handshake of an outbound dial. Shorter frees a stuck handshake's slot sooner, at the risk of a false timeout under real load.
max_pending_handshakesinteger256Cap on concurrently in-progress inbound handshake attempts. Higher rides out a bigger connection burst, at the cost of more concurrent in-flight handshake buffers and tasks.
max_frame_sizebytes"16MB"Cap on a single wire frame's payload: a handshake, a chunk, or a message smaller than a chunk. A larger message travels as chunks, up to its traffic lane's queue size. [cluster.consensus] max_append_bytes may be at most half of the default.
tcp_nodelaybooltrueSets TCP_NODELAY on every peer connection, dialed or accepted (under TLS too), so a small frame leaves at once instead of waiting on Nagle's algorithm and the peer's delayed acknowledgement.
tcp_user_timeoutduration"2000ms"Linux TCP_USER_TIMEOUT: how long sent data may stay unacknowledged, or queued behind a peer that stopped reading, before the kernel closes the connection. "0" leaves the kernel default (minutes).
keepalive_idleduration"1s"TCP keepalive: idle time before the first probe. The kernel counts whole seconds, so a sub-second value rounds up. "0" turns keepalive off.
keepalive_intervalduration"500ms"TCP keepalive: time between probes, applied as 1 s by the kernel's whole-second resolution. On Linux tcp_user_timeout bounds detection anyway.
keepalive_probesinteger2TCP keepalive: unanswered probes before the kernel drops the connection. 1 to 127.
ping_intervalduration"500ms"A connection that has sent nothing for this long sends a liveness ping, so its peer always has something to hear; a busy connection never pings. "0" turns pings off, which requires dead_peer_timeout = "0" too.
dead_peer_timeout"auto" or duration"auto"A connection on which no byte arrives from the peer for this long is closed and reported lost. This catches a peer whose kernel still answers while its process is stalled, which neither the user timeout nor keepalive can see. "auto" derives it per peer from the measured ping round trip (two ping intervals plus twice the round trip's rarely exceeded bound), inside dead_peer_timeout_min and dead_peer_timeout_max; a duration pins one value for every peer; "0" turns it off. It must be longer than twice ping_interval.
dead_peer_timeout_minduration"2s"Floor of the adaptive deadline, so the jitter of a long link is never read as a dead peer.
dead_peer_timeout_maxduration"10s"Ceiling of the adaptive deadline: the longest a stalled peer can go unnoticed, even on the worst link.
peer_event_bufferinteger1024Capacity of the internal channel that tells each subsystem a peer connection was established or lost. A subsystem that falls this far behind treats every peer as lost once, which is always safe. Counted in the process memory budget at 512 bytes per slot.
chunk_sizebytes"64KB"Messages larger than this travel between nodes as chunks, and each traffic lane's connection takes its streams in turn a chunk at a time, so a large transfer delays a small message on the same connection by at most one chunk (0.5 ms at 1 Gbit/s). Between 4KB and max_frame_size.
control_queue_size"auto" or bytes"auto"The most bytes of control traffic (membership, consensus, the commit protocol, forwarded DDL and lock calls, shuffle fetches) this node holds queued to send to all peers, and separately received and not yet consumed. "auto" is half of [cluster] recv_queue_bytes, at least one message of max_frame_size.
interactive_queue_size"auto" or bytes"auto"The same for interactive traffic (replies to forwarded point reads). "auto" is a quarter of recv_queue_bytes, at least one message of max_frame_size.
bulk_queue_size"auto" or bytes"auto"The same for bulk traffic (query data exchange, fragment dispatch with its broadcast join side, snapshot transfer). "auto" is recv_queue_bytes, at least the largest broadcast join side (128MB) plus one message of max_frame_size. No message larger than its lane's queue is sent.
bulk_send_bufferbytes"4MB"The kernel send buffer (SO_SNDBUF) of each bulk connection. Larger keeps a longer link full, at the cost of kernel memory per peer.
bulk_unsent_limit"auto" or bytes"auto" (two chunks)The most bytes a bulk connection's kernel keeps queued and not yet sent (TCP_NOTSENT_LOWAT, Linux and macOS), so the rest waits in the node's own queue, where every stream gets its turn.

Every stream of a lane gets its own window of the lane (the lane's size shared by the peers), so a subsystem that stops reading holds only its own window and stops only its own senders; every other stream keeps flowing. A lane's connection lost to a peer is reported for that lane's traffic only.

Each connection's deadline in force, and whether it sits on the floor, on the ceiling, in between or is pinned, is exported on /metrics (see Observability). These timers never decide a correct answer. They only turn a silent hang into an explicit connection loss: a replica is re-probed, and a transaction step forwarded to the lost node is reported in doubt at once instead of waiting out its own deadline.

[cluster.distributed_join]​

How a join whose two sides live on different nodes adapts while it runs: merging small transfers, splitting a skewed partition, retrying and backing up a slow consumer, and the size of the filter a broadcast join ships. Formerly the top-level [aqe] table (see Renamed keys). The coalescing, skew-split and straggler behaviors only run once a session opts into them; see Distributed queries.

KeyTypeDefaultEffect
coalesce_targetbytes/string, ≥ 1"64MB"Target bytes for one coalesced shuffle-join consumer read.
skew_factorinteger, ≥ 15A partition is flagged as skewed when its bytes exceed this multiple of the median.
skew_minbytes/string, ≥ coalesce_target"256MB"Absolute floor a partition must also exceed to be flagged as skewed.
consumer_attemptsinteger, ≥ 13Bounded re-dispatch attempts for a materialized consumer fragment before the query fails.
straggler_backup_delayduration, > 0"2s"How long a fragment may run before a speculative backup is dispatched on another node.
runtime_filter_max_keysinteger, ≥ 1262144Broadcast-join small-side row budget past which the runtime Bloom filter is skipped.

[cluster.exchange]​

How a node receives the rows other nodes send it during a distributed query, and the threads that run that work. Formerly exchange_recv_timeout and exchange_channel_capacity under [execution] (see Renamed keys).

KeyTypeDefaultEffect
receive_timeoutduration, > 0"30s"How long a query step that receives rows streamed from other nodes waits for all of its senders to finish before it fails the query. It bounds a wedged or unreachable sender, not a healthy transfer.
receive_buffer_chunksinteger, ≥ 132Row batches one receiving endpoint buffers before its senders wait. A deeper buffer smooths a rate mismatch between senders and receiver, at the cost of more memory per endpoint.
threadsintegera quarter of the cores, 2 to 8Threads that run the distributed query work other nodes send this node (their fragments, the rows arriving and the results sent back), off the consensus threads. Absent, they are sized from the cores like the consensus threads. 1 to 256.

[cluster.log_compaction]​

When a group compacts its replication log behind a snapshot of its state, how far a lagging replica may fall behind before it is repaired from that snapshot instead, and how much memory snapshots may hold. Every entry count adapts to the group's own snapshot and entry sizes when absent; setting it pins it.

KeyTypeDefaultEffect
trigger_entriesintegeradaptiveRetained log entries above which a data group compacts its log behind a snapshot of its state. Adaptive: about as many entries as fill the group's own snapshot, between 1,000 and 1,000,000. At least 1.
retained_entriesintegeradaptive, a tenth of the triggerEntries kept below the compaction point, so a replica a little behind catches up by replication instead of a snapshot. Must be below trigger_entries.
replica_lag_limit_entriesintegeradaptive, equal to the triggerA replica further behind stops holding the leader's log and is repaired from a snapshot instead; one still catching up gets twice this. So a dead replica bounds the log instead of growing it. At least 1.
metadata_trigger_entriesintegeradaptiveRetained entries above which the cluster metadata log compacts. At least 1.
memory_bytesbytestwo apply mailboxes (8 MiB each) per thread of the node's cluster share (a quarter of the cores, 2 to 8): 32MB to 128MBMemory building, sending and receiving state snapshots may hold at once. A snapshot that would exceed it is given up and retried later, never granted over it. Counted in the process memory budget.

[cluster.apply]​

How a node applies committed replicated writes. A committed entry is decided first (the waiting session gets its answer), then its rows are written on dedicated apply threads with one data log sync per batch; reads wait until a row is applied and synced.

KeyTypeDefaultEffect
reserve_io_coresbooltrueKeep a quarter of this node's cores (at most eight) for consensus, replication and client connections, so heavy queries cannot delay them; the query workers take the rest. false shares every core with queries, as a single node does. A node with fewer than four cores reserves none. When several nodes share one machine without separate CPU sets, they reserve the same cores; give each its own CPU set or set this to false.
threadsintegera quarter of the cores, 2 to 8Threads that apply committed writes to this node's storage, off the consensus threads. 1 to 256.

[cluster.dilith]​

The single-round commit protocol of distributed transactions. decision_window, retention_window and fence_ahead_factor are parameters of its correctness argument and are fixed values; every other key only times when something is looked at or re-sent, never what is decided, and no wall clock decides an outcome. An absent table is every key at its default.

KeyTypeDefaultEffect
decision_windowinteger256Decisions a group remembers beyond the commits it still owes a discharge; a vote or sweep older than this is refused. At least 1.
retention_windowinteger64Positions the retention floor trails a group's closed position; a read below the floor is refused and retried. At least 1.
fence_ahead_factorinteger4How many recent gaps between analytical cuts a fence sets a never-written group's floor ahead, so later queries under it need no entry.
recovery_timeoutduration"500ms"Wait behind an undecided transaction before its recovery starts.
pending_vote_recovery_factorinteger10A group leader recovers a vote pending this many recovery timeouts with nobody waiting behind it. At least 1.
recent_verdicts_keptinteger256Recovery verdicts a node remembers, to resend a known outcome instead of probing again. At least 1.
probe_retryduration"360ms"First re-ask interval of a recovery or status probe.
probe_retry_growthfloat2.0Factor each unanswered probe round multiplies the interval by. At least 1.
probe_retry_stepsinteger4Rounds after which a probe's interval stops growing.
horizon_retryduration"90ms"Re-ask interval of a watermark round: an analytical query's, and a REPEATABLE READ or SERIALIZABLE snapshot's while every replica of a shard group refuses it.
max_forward_hopsinteger2Times a leader-bound request may be forwarded toward the current leader.
parked_retry_initialduration"2ms"First re-check interval of a request parked at a group.
parked_retry_maxduration"64ms"Ceiling the parked re-check interval doubles up to while it finds nothing to do.
leaderless_retry_maxduration"4ms"Re-check ceiling while a request is held for a group with no leader.
leaderless_holdduration"2500ms"How long a replica holds a vote or outcome for a group with no leader before dropping it to the sender's retry.
held_notice_recheckduration"300ms"How often a coordinator re-checks a request a replica said it is holding.
younger_waitduration"1500us"How long the younger of two conflicting transactions may wait before it is refused.
younger_wait_queue_maxinteger4Parked votes at a group beyond which a younger transaction is refused at once.
conflict_count_decayduration"200ms"Decay time of a group's conflict count.
conflict_hot_thresholdfloat6.0Decayed conflict count at or above which a younger transaction is refused without waiting.
refusal_count_decayduration"500ms"Decay time of the refusal counts that decide which groups are asked to vote first.
refusing_thresholdfloat1.0Decayed refusal count at or above which a group counts as refusing.
ask_first_thresholdfloat0.15Decayed refusal count at or above which a group is asked to vote before the rest. At most refusing_threshold.
refusal_counts_maxinteger64Refusal counts a coordinator keeps before it starts over. At least 1.
refusing_groups_trackedinteger8Groups a coordinator tracks as refusing somebody. At least 1.
outcome_flush_delayduration"0ms"How long a leader holds a decided outcome before appending it alone.
outcome_ack_graceduration"40ms"Extra wait, beyond two round trips, before an outcome is re-sent.
discharge_batch_delayduration"25ms"How long a coordinator batches discharge notices per destination.
discharge_carry_waitduration"250ms"How long a leader lets discharge notices wait for an entry to carry them.
round_trip_initialduration"30ms"Round-trip estimate for a node before its first answer is measured.
round_trip_smoothingfloat0.25Weight of each new round-trip sample in the running estimate; above 0, at most 1.
round_trip_sample_maxduration"5s"Round-trip samples above this are ignored.
retry_round_tripsfloat3.0A request (a commit's vote, or a round of a COPY's staged rows) is first re-sent after this many measured round trips to its destination; a round of rows counts its bytes at the rate the node's rounds have moved theirs in its round trip. At least 1, so a re-send never precedes its answer.
retry_minduration"2ms"Shortest re-send interval.
retry_maxduration"1200ms"Longest re-send interval, unless one round trip is longer.
retry_growthfloat4.0Factor each re-send of the same request multiplies its interval by. At least 1.
retry_growth_stepsinteger8Re-sends after which the interval stops growing.
heartbeat_align_intervalduration"1s"Period of the one-byte entry a node appends to the groups it leads so their heartbeats travel together; stretches by the square root of the fan-out above heartbeat_align_reference_groups.
heartbeat_align_min_groupsinteger2Groups a node must lead before it aligns their heartbeats at all.
heartbeat_align_reference_groupsfloat32.0Groups led per peer above which the alignment period stretches.
copy_stage_idle_timeoutduration, optional, greater than zeroten times [cluster.consensus] election_timeout_maxHow long a group keeps the staged rows of a COPY whose node stopped sending or renewing them before the group's leader drops them; the sending node waits at most half of it for a round's answer before it sends the round again. Only decides when: rows a commit's vote holds are never dropped.
owed_commit_retentionduration, optional, at least [cluster.transactions] commit_answer_timeout + outcome_query_timeoutthat sum plus recovery_timeout times pending_vote_recovery_factor (16s at the defaults)How long a group keeps a committed transaction whose last notice was lost (its node stopped right after answering its client) before the group's leader finishes it itself. Only decides when: no commit's outcome ever changes.
protocol_state_memorybytes/string, ≥ 1, optionala 32nd of the machine's memory, 32MB to 4GBMemory the commit protocol's tables may hold on this node: pending votes, commits not yet discharged, what the decision and retention windows keep, the outcomes a leader carries and the room reserved for votes being appended. A group leader appends a new vote only within it; a vote that finds no room waits, oldest transaction first, until the protocol's own steps free some, and is never refused for memory. Counted in the process memory budget. Only decides when.

A complete, working example that sets [cluster]'s core fields and some of these sub-tables at their defaults ships inside the Docker image itself; see Start from the shipped example for the exact path, how to pull it out and edit it, and a pointer to a larger annotated example that does cover them.

[execution]​

KeyTypeDefaultEffect
workers"auto" or a positive integer"auto" (every core the process can see)Core-pinned workers a query's morsel scheduling dispatches across. An explicit count caps dispatch. 0 or a negative value is a hard parse-time rejection, not a silent clamp to 1.
morsel_sizebytes/string"4MB"Target bytes per morsel for byte-size-aware table scans.
transactional_morsel_thresholdusize, 1-644HTAP dispatch classifier: a statement dispatching this many morsels or fewer, or any DML statement, classifies as transactional and rides the scheduler's priority lane, dequeued ahead of analytical work, so a short OLTP-style statement no longer waits behind a whole analytical query. An out-of-range value is a hard startup error naming the accepted range, never a silent clamp.
aging_promotion_msu64 (milliseconds), 1-100010HTAP starvation ceiling: analytical work waiting longer than this is served ahead of further transactional work, so a starved analytical task is guaranteed service within this window even under a sustained flood of short statements. Same startup validation as transactional_morsel_threshold.
copy_batch_rowsusize65536Rows buffered per COPY FROM batch before insertion.
copy_pipeline_depthusize, ≥ 14Parsed COPY batches buffered ahead of insertion. It also bounds COPY ... TO STDOUT: that many reused buffers of about 256 KiB each per export, whatever its size.
max_recursive_iterationsusize1000Fixpoint iteration cap for a recursive CTE before it's treated as non-terminating.
spill_agg_partitionsusize, ≥ 164Partitions for spill-to-disk hash aggregation.
grace_join_min_partitionsusize, ≥ 1 and ≤ grace_join_max_partitions16Floor of the adaptive Grace hash join partition count.
grace_join_max_partitionsusize256Ceiling of the adaptive Grace hash join partition count.
join_bloom_thresholdusize, ≥ 11024Minimum build-side rows before a hash join builds a Bloom pre-filter.
parallel_build_thresholdusize10000Minimum build-side rows before a hash join build parallelizes across threads.
runtime_filtersbooltrueEnables sideways information passing: a selective hash-join build side publishes a Bloom filter that prunes non-matching probe-side rows, and its key range as two live zone-map bounds that let the probe-side scan skip pages that cannot join before any I/O. Disabling it removes both prunings; the filters are always advisory, so results are unchanged either way.
arena_chunk_sizebytes/string, ≥ 1"2MB"Chunk size for the per-query bump allocator.
sequence_reservation_blockusize32Values a sequence durably reserves ahead of use; a restart may skip up to this many values.
jit_prefetch_distanceu648Rows ahead the compiled kernels software-prefetch hash-table buckets. The effective distance is additionally clamped by a fill-buffer budget model, so a kernel prefetching several streams at once cannot stall the core by exhausting its L1 fill buffers; raising this past what the budget allows changes nothing. 0 disables prefetch emission entirely. Read at code-generation time, so it applies to kernels compiled after the change.
jit_compile_threadsusize, clamped to 1-83Size of the background compile pool. At the default, the tier ladders of different queries compile in parallel instead of queueing behind one worker, which is what keeps a freshly seen query shape from running interpreted while a backlog drains. Each worker carries a 16 MiB stack and compiles are CPU-bound, so a large pool just steals cores from query workers. Set 1 to restore strict single-threaded compilation as a rollback lever.
strict_compiledboolfalseAccepted so existing configurations keep loading; it no longer changes anything. Once a statement's compiled code is ready, every chunk runs compiled, and a function the compiled code has no entry for is an error, whatever this key says.
jit_alias_factsboolfalseEmit noalias facts into the generated IR.
jit_profile_factsboolfalseFeed observed selectivity into the tier-2 compiler as branch expectations.
jit_prefer_vector_widthu640 (no attribute)Sets prefer-vector-width on generated kernels.
copy_consumersusize0 (auto)Consumer threads a bulk load uses. 0 derives it from the pool width, clamped to 1-8; an explicit 1 restores today's single async consumer (the kill switch). Only has an effect on a file-path CSV/TSV COPY with no column list, into a table with no indexes, no partitioning, no declared foreign key, and no distributed routing - every other shape takes the single consumer regardless of this value.
copy_parse_lane_percentu8, 0-1000 (auto)Percent of the pool width used as CSV parse lanes during a load. 0 picks the measured default.

Renamed keys​

A few keys have moved to a new name. The old name keeps working in every future release, so a file written for an earlier release starts unchanged; a node started with an old name says so once in its startup log, naming the new one. A file that sets both the old and the new name of one key is refused at startup, naming both.

Old nameNew name
[cluster.raft][cluster.consensus]
[aqe][cluster.distributed_join]
[execution] exchange_recv_timeout[cluster.exchange] receive_timeout
[execution] exchange_channel_capacity[cluster.exchange] receive_buffer_chunks
[vector] distributed_two_phase_bytes[cluster.vector] two_phase_read_bytes

The last four only mean something on a cluster. In a file with no [cluster] table they have no effect, and the startup log says so; their values are still checked, so a value refused on a cluster is refused there too.

Retired keys​

A few cluster keys tuned the previous commit path, which the commit protocol replaced; the commit protocol does what each did without a setting. A file that still sets one starts unchanged: the key is read by nothing, never refused, and the startup log says once that it has no effect.

KeyWhy it has no effect
[cluster] commit_protocol, [cluster] commit_v2Every cluster transaction commits through the one commit protocol.
[cluster] tso_batchThe commit protocol orders commits without a timestamp service.
[cluster] txn_lock_ttlTransactions on cluster tables hold no row lock leases, and the commit protocol recovers an abandoned commit itself ([cluster.dilith] recovery_timeout).
[cluster.txn]The commit protocol keeps the history it validates reads against itself ([cluster.dilith] retention_window).
[cluster.apply] outcome_retention_max_entriesThe commit protocol answers each commit itself, so no table of outcomes is kept.

[statistics]​

Automatic statistics collection for the query planner.

KeyTypeDefaultEffect
auto_analyze_check_intervalduration, > 0"120s"Interval between auto-analyze cycles.
auto_analyze_staleness_ratiof64, > 00.10Re-analyze a table once its row count has drifted by more than this fraction since the last analysis.
auto_analyze_sample_pagesusize, ≥ 164Max pages sampled per table when building histograms and most-common-value lists.
mcv_entriesusize, ≥ 1100Max most-common-value entries retained per column.
column_pairsusize, ≥ 150Max column pairs tracked for multi-column statistics per table.
analyze_on_loadbooltrueTrigger a full ANALYZE once a COPY or bulk INSERT commits, instead of waiting for the next staleness sweep.

[optimizer]​

Join-ordering search limits and the statement plan cache.

KeyTypeDefaultEffect
dp_table_limitusize, ≥ 112Base relation count at or below which join ordering runs the exact dynamic-programming search; above it, a greedy fallback orders the join instead.
cascades_budget_factoru64, ≥ 1500Cascades iteration budget multiplier: max_iterations = num_tables^2 * this.
semi_reversal_skew_factorf64, ≥ 1.04.0Skew factor above which a semi-join is considered for reversal. Only consulted when the SCRAMDB_T4_SEMI_REVERSE environment variable is set to 1.
plan_cache_memorysize (bytes or "64MB")unset: 1/256 of machine memory, between 8 MiB and 512 MiBMemory the statement plan cache may hold (plans, their compiled code and plan markers). A value pins it.

[oltp]​

KeyTypeDefaultEffect
point_fast_pathbooltrueClassifies and inline-executes an eligible single-table index-equality SELECT / UPDATE / DELETE, or a single-row literal INSERT, bypassing planning, bytecode compilation, and the pipeline scheduler. A pure latency optimization: turning it off, or a statement that doesn't classify, changes nothing about the result.

[resilience]​

Client liveness checks, statement and session timeouts, and the connection-pool watchdog.

KeyTypeDefaultEffect
client_liveness_check_intervalduration"2s"Cadence of the non-blocking peer-liveness probe. "0" disables the probe entirely; that's an intentional off switch, not an error.
statement_timeoutduration"0" (disabled)Server-wide default: cancels a running statement past this duration. SET statement_timeout = ... overrides it per-session, live, with no restart. Matches PostgreSQL's own default of off.
idle_in_transaction_session_timeoutduration"0" (disabled)Server-wide default: rolls back and closes a session that's been idle inside an open transaction (including an aborted one) past this duration. SET idle_in_transaction_session_timeout = ... overrides it per session. Matches PostgreSQL's own default.
transaction_timeoutduration"0" (disabled)Server-wide default: rolls back and closes any one transaction, running or idle, once it's been open this long. SET transaction_timeout = ... overrides it per session. Matches PostgreSQL 17's transaction_timeout.
lock_timeoutduration"30s"How long a statement waits for a lock before failing. PostgreSQL's own default for this setting is "0" (wait forever); ScramDB deliberately does not do that.
deadlock_timeoutduration"1s"How long to wait before running deadlock detection.
max_prepared_transactionsusize0Prepared transactions allowed at once. 0 disables two-phase commit.
prepared_transaction_timeoutduration"3600s"How long a prepared transaction may stay prepared before it is rolled back.
tcp_keepalivebooltrueOS-level TCP keepalives on every accepted pgwire socket.
tcp_keepalive_idleduration"60s"Idle time before the OS sends the first TCP keepalive probe.
enabledbooltrueMaster switch for this section's connection-pool watchdog.
watchdog_enabledbooltrueKill switch for just the pool watchdog's observer thread.
cadence_floorduration, > 0"1ms"Fastest pool-watchdog sampling cadence.
cadence_capduration, ≥ cadence_floor"10ms"Slowest pool-watchdog sampling cadence.
  • Configuration Reference - how these sections and the other two config surfaces fit together.
  • Environment Variables - the canonical list of env vars that exist outside this file.
  • Precedence - which of the config file, a CLI flag, or an env var wins for the handful of settings more than one of them can touch.