Skip to main content

Vector index configuration

By the end of this page you will know every [vector] config-file key (the server-wide defaults for anode/manode indexes), how to change one live with ALTER SYSTEM, and where the docker and cluster config files set them. For the SQL surface these settings back (CREATE INDEX ... WITH (...), ALTER INDEX, session parameters), see Statements and Vector Search; for what each knob trades off, see Tuning: Vector Indexes.

[vector]​

Absent entirely, byte-for-byte identical to writing the table out with every field at its default below, the same rule every [section] in Config File follows.

Liveness says when a changed value takes hold: live applies to the next statement that reads it; next regrow applies the next time an index's arena grows to hold more vectors; next open applies the next time an index opens (a fresh build, or REINDEX).

KeyTypeDefaultLivenessEffect
memory_budget_totalpercentage or byte size"25%" (of the node's resolved execution memory budget, [storage] execution_memory_bytes/execution_memory_percent)liveThe ceiling for every vector index arena on this node together. Raising it applies at once; lowering it below what is already reserved is refused, naming the reserved amount
index_memory_budgetbytes"1GB"next regrowDefault per-index ceiling, overridden per index by CREATE INDEX ... WITH (memory_budget = ...)
arena_slotsinteger, optionalunset (automatic)next regrowThe arena's promised vector capacity; unset computes max(live_rows * arena_growth_factor, arena_min_slots) clamped to the ceiling. WITH (arena_slots = n) overrides per index
arena_growth_factorinteger, ≥ 14liveMultiplier on live row count used to size a fresh arena promise
arena_min_slotsinteger, ≥ 11048576liveFloor on the arena promise regardless of live row count
arena_slack_percentinteger, percent5liveSlack over the live row count applied when sizing a REINDEX
probe_permilleinteger, 0-1000100liveANODE's own default; parts per thousand of the coarse cells a search probes
near_permilleinteger, 0-1000250liveParts per thousand of the probed cells nearest the query, weighted more heavily
walk_budgetinteger, ≥ 1900liveGraph-walk candidate budget
beaminteger, ≥ 164liveGraph-walk beam width
rerank_extrainteger, ≥ 08liveExtra candidates carried into the exact re-rank beyond what k and over_fetch alone would ask for
rank_headinteger, ≥ 1256liveRanked coarse cells a search may visit
over_fetchinteger, ≥ 12liveCandidates fetched per requested row before the exact re-rank
rerank_chunkinteger, ≥ 11024liveCandidates per batched row fetch and kernel call in the exact re-rank
max_scan_tuplesinteger, ≥ 120000liveThe widening and threshold-search candidate cap
iterative_scan"off" | "relaxed_order" | "strict_order""strict_order"liveWhether an underfilled indexed search widens, and whether it re-sorts across rounds; see Vector Search
exact_tier_ratiointeger, ≥ 14liveTier 1 (exact scan, no index call) is chosen at or under this many expected matches per requested row
selectivity_boundaryfloat, 0.0-1.00.55liveTier 2 (allow-list traversal) up to this predicate selectivity; tier 3 (post-filter widening) above it
threshold_slackfloat, ≥ 0.00.10liveSlack over the requested bound a distance-threshold search's walk may exceed before stopping
knn_join_batchinteger, ≥ 164liveQuery vectors per worker chunk in a nearest-neighbour join
compact_after_bytesbytes"1GB"liveJournal bytes written since the last compaction past which a background compaction triggers
compact_tombstone_ratiofloat, 0.0-1.00.2liveTombstones over live vectors past which a background compaction triggers
compaction_check_interval_msinteger, ≥ 1 (ms)30000liveHow often the background lane checks every index against the compaction triggers. Three events do not wait for it: an index that runs out of arena, a raise of memory_budget_total, and an index option change each run the lane's next pass at once
flush_interval_secsinteger, ≥ 1 (seconds)30liveJournal flush cadence between checkpoints
maintenance_workersinteger0 (half the visible cores, at most 8)live (next job)Concurrent ANODE fits and backfills this node runs at once, shared across every index by the maintenance gate
build_workersinteger0 (the execution worker pool's width)next buildANODE's parallel link width for a build
rerank_prefetchinteger, ≥ 18next openANODE's rerank prefetch depth
simd_level"auto" | "scalar" | "avx2" | "avx512-256" | "avx512" | "neon" | "neon+dotprod""auto"next openThe instruction set an index opens with; only ever a step down from what the CPU actually has
sparsevec_index_max_dimsinteger, ≥ 12000liveThe most dimensions a SPARSEVEC column may have for CREATE INDEX (pgvector's limit); a wider column is refused. An index built under a higher value keeps working after it is lowered.
ledger_tail_rowsinteger, ≥ 11048576liveRows the row-id ledger's reverse-lookup tail map holds before a background rebuild of its sorted array
sample_min_interval_msinteger, ≥ 1 (ms)250liveFloor between observability samples, so a scrape storm cannot become an ANODE metrics storm

Cluster settings: [cluster.vector]​

The vector settings that only mean something on a cluster live under the [cluster] table. Each is also a live setting under the runtime name shown, for SHOW and ALTER SYSTEM SET.

KeyTypeDefaultRuntime nameEffect
recall_enabledbooleantruevector.recall_enabledWhether each node measures the recall of the manode index parts it hosts (scram_vector_index_parts.recall_at_10); false turns the job off
recall_intervalduration, up to one week"0" (adaptive)vector.recall_interval_secsWait before a changed part's recall is measured again. At "0" the wait follows the part's churn and the measurement's own cost: a part nothing was written to is never measured again, a changed one waits its cost floor divided by the share of it rewritten, between that floor (100 times the CPU its last measurement took, so the job holds to one percent of one core) and one hour. A duration pins the wait
recall_probesinteger, 0-10240 (adaptive)vector.recall_probesProbe queries one recall measurement samples. At 0: enough for a one-point standard error at the part's last recall, between 8 and 64, and never more than the vector memory budget can hold. A positive value pins it
recall_max_rowsinteger, ≥ 11000000vector.recall_max_rowsThe largest part, in live vectors, a recall measurement reads exactly; a larger part is not measured and says so in recall_note
stats_node_timeoutduration, 100ms to 10 minutes"5s"vector.stats_node_timeout_msHow long a whole-cluster vector statistics view waits for one node before it reports that node as unreachable
two_phase_read_bytesbytes"256KB"vector.distributed_two_phase_bytesAbove this estimate of the bytes a one-phase read across nodes would ship, the read uses the two-phase mode instead (ids and distances first, rows second). Formerly [vector] distributed_two_phase_bytes, a name that keeps working (see Renamed keys)

Every byte-size field accepts the same grammar as the rest of the config file (a plain integer, or "4GB"/"512MB"/etc; see Config File: File format). An unrecognized key under [vector], or a value outside the ranges above, is a hard parse or validation error naming the exact field, never a silently clamped or ignored setting.

[vector]
memory_budget_total = "8GB"
index_memory_budget = "512MB"
over_fetch = 4
simd_level = "avx2"

ALTER SYSTEM SET vector.<key>​

Every key above (and every per-index search knob) is also a live runtime setting, changeable without a restart:

ALTER SYSTEM SET vector.over_fetch = 4;
ALTER SYSTEM RESET vector.over_fetch;
SELECT pg_reload_conf();

ALTER SYSTEM in this release accepts vector.<key> only (superuser required); every other section is a named follow-up. Three keys take a thousandths integer here instead of the config file's plain fraction, since a runtime setting cell is always an integer:

Config-file key (fraction)ALTER SYSTEM / per-index key (integer, 0-1000)
selectivity_boundaryvector.selectivity_boundary_permille
threshold_slackvector.threshold_slack_permille
compact_tombstone_ratiovector.compact_tombstone_ratio_permille

Every other key keeps its config-file name and form, vector.memory_budget_total included (it accepts the same "25%" or "8GB" text ALTER SYSTEM SET gives it as a config file does).

Precedence, highest first: a session parameter (SET anode.<key> = ..., statement scope) > an index option (ALTER INDEX ... SET) > a catalog override (ALTER SYSTEM SET) > this config file > the built-in default. SHOW vector.<key> and the scram_vector_settings system view (see Observability) report both the value in force and which of those five layers set it.

Node-scoped keys: memory_budget_total, maintenance_workers, build_workers, rerank_prefetch, simd_level, and sample_min_interval_ms are stored per node in a cluster rather than cluster-wide. ALTER SYSTEM SET vector.<key> = <value> ON NODE '<id>' targets one node; the plain form (no ON NODE) targets the issuing node for these six keys, and every node for the rest. Every other section of ALTER SYSTEM is refused in this release, naming the reason.

Docker and cluster config files​

The [vector] table follows the same "absent = every field at its default" rule in every config file ScramDB ships or reads: docker/default-config.toml, docker/cluster-config.template.toml (rendered by the container entrypoint), cluster-config.example.toml, local-config.toml, and the Kubernetes ConfigMap that feeds a Helm-deployed node's config file. Set only the keys you actually want to override; the rest inherit their built-in defaults exactly as an omitted [vector] table would. See Clustering and Kubernetes for where each of those files lives in a real deployment.