Vector index configuration
By the end of this page you will know every [vector] config-file key (the server-wide
defaults for anode/manode indexes), how to change one live with ALTER SYSTEM, and
where the docker and cluster config files set them. For the SQL surface these settings
back (CREATE INDEX ... WITH (...), ALTER INDEX, session parameters), see
Statements
and Vector Search; for what each knob trades off,
see Tuning: Vector Indexes.
[vector]
Absent entirely, byte-for-byte identical to writing the table out with every field at
its default below, the same rule every [section] in Config File
follows.
Liveness says when a changed value takes hold: live applies to the next
statement that reads it; next regrow applies the next time an index's arena grows
to hold more vectors; next open applies the next time an index opens (a fresh
build, or REINDEX).
| Key | Type | Default | Liveness | Effect |
|---|---|---|---|---|
memory_budget_total | percentage or byte size | "25%" (of the node's resolved execution memory budget, [storage] execution_memory_bytes/execution_memory_percent) | live | The ceiling for every vector index arena on this node together. Raising it applies at once; lowering it below what is already reserved is refused, naming the reserved amount |
index_memory_budget | bytes | "1GB" | next regrow | Default per-index ceiling, overridden per index by CREATE INDEX ... WITH (memory_budget = ...) |
arena_slots | integer, optional | unset (automatic) | next regrow | The arena's promised vector capacity; unset computes max(live_rows * arena_growth_factor, arena_min_slots) clamped to the ceiling. WITH (arena_slots = n) overrides per index |
arena_growth_factor | integer, ≥ 1 | 4 | live | Multiplier on live row count used to size a fresh arena promise |
arena_min_slots | integer, ≥ 1 | 1048576 | live | Floor on the arena promise regardless of live row count |
arena_slack_percent | integer, percent | 5 | live | Slack over the live row count applied when sizing a REINDEX |
probe_permille | integer, 0-1000 | 100 | live | ANODE's own default; parts per thousand of the coarse cells a search probes |
near_permille | integer, 0-1000 | 250 | live | Parts per thousand of the probed cells nearest the query, weighted more heavily |
walk_budget | integer, ≥ 1 | 900 | live | Graph-walk candidate budget |
beam | integer, ≥ 1 | 64 | live | Graph-walk beam width |
rerank_extra | integer, ≥ 0 | 8 | live | Extra candidates carried into the exact re-rank beyond what k and over_fetch alone would ask for |
rank_head | integer, ≥ 1 | 256 | live | Ranked coarse cells a search may visit |
over_fetch | integer, ≥ 1 | 2 | live | Candidates fetched per requested row before the exact re-rank |
rerank_chunk | integer, ≥ 1 | 1024 | live | Candidates per batched row fetch and kernel call in the exact re-rank |
max_scan_tuples | integer, ≥ 1 | 20000 | live | The widening and threshold-search candidate cap |
iterative_scan | "off" | "relaxed_order" | "strict_order" | "strict_order" | live | Whether an underfilled indexed search widens, and whether it re-sorts across rounds; see Vector Search |
exact_tier_ratio | integer, ≥ 1 | 4 | live | Tier 1 (exact scan, no index call) is chosen at or under this many expected matches per requested row |
selectivity_boundary | float, 0.0-1.0 | 0.55 | live | Tier 2 (allow-list traversal) up to this predicate selectivity; tier 3 (post-filter widening) above it |
threshold_slack | float, ≥ 0.0 | 0.10 | live | Slack over the requested bound a distance-threshold search's walk may exceed before stopping |
knn_join_batch | integer, ≥ 1 | 64 | live | Query vectors per worker chunk in a nearest-neighbour join |
compact_after_bytes | bytes | "1GB" | live | Journal bytes written since the last compaction past which a background compaction triggers |
compact_tombstone_ratio | float, 0.0-1.0 | 0.2 | live | Tombstones over live vectors past which a background compaction triggers |
compaction_check_interval_ms | integer, ≥ 1 (ms) | 30000 | live | How often the background lane checks every index against the compaction triggers. Three events do not wait for it: an index that runs out of arena, a raise of memory_budget_total, and an index option change each run the lane's next pass at once |
flush_interval_secs | integer, ≥ 1 (seconds) | 30 | live | Journal flush cadence between checkpoints |
maintenance_workers | integer | 0 (half the visible cores, at most 8) | live (next job) | Concurrent ANODE fits and backfills this node runs at once, shared across every index by the maintenance gate |
build_workers | integer | 0 (the execution worker pool's width) | next build | ANODE's parallel link width for a build |
rerank_prefetch | integer, ≥ 1 | 8 | next open | ANODE's rerank prefetch depth |
simd_level | "auto" | "scalar" | "avx2" | "avx512-256" | "avx512" | "neon" | "neon+dotprod" | "auto" | next open | The instruction set an index opens with; only ever a step down from what the CPU actually has |
sparsevec_index_max_dims | integer, ≥ 1 | 2000 | live | The most dimensions a SPARSEVEC column may have for CREATE INDEX (pgvector's limit); a wider column is refused. An index built under a higher value keeps working after it is lowered. |
ledger_tail_rows | integer, ≥ 1 | 1048576 | live | Rows the row-id ledger's reverse-lookup tail map holds before a background rebuild of its sorted array |
sample_min_interval_ms | integer, ≥ 1 (ms) | 250 | live | Floor between observability samples, so a scrape storm cannot become an ANODE metrics storm |
Cluster settings: [cluster.vector]
The vector settings that only mean something on a cluster live under the [cluster] table.
Each is also a live setting under the runtime name shown, for SHOW and ALTER SYSTEM SET.
| Key | Type | Default | Runtime name | Effect |
|---|---|---|---|---|
recall_enabled | boolean | true | vector.recall_enabled | Whether each node measures the recall of the manode index parts it hosts (scram_vector_index_parts.recall_at_10); false turns the job off |
recall_interval | duration, up to one week | "0" (adaptive) | vector.recall_interval_secs | Wait before a changed part's recall is measured again. At "0" the wait follows the part's churn and the measurement's own cost: a part nothing was written to is never measured again, a changed one waits its cost floor divided by the share of it rewritten, between that floor (100 times the CPU its last measurement took, so the job holds to one percent of one core) and one hour. A duration pins the wait |
recall_probes | integer, 0-1024 | 0 (adaptive) | vector.recall_probes | Probe queries one recall measurement samples. At 0: enough for a one-point standard error at the part's last recall, between 8 and 64, and never more than the vector memory budget can hold. A positive value pins it |
recall_max_rows | integer, ≥ 1 | 1000000 | vector.recall_max_rows | The largest part, in live vectors, a recall measurement reads exactly; a larger part is not measured and says so in recall_note |
stats_node_timeout | duration, 100ms to 10 minutes | "5s" | vector.stats_node_timeout_ms | How long a whole-cluster vector statistics view waits for one node before it reports that node as unreachable |
two_phase_read_bytes | bytes | "256KB" | vector.distributed_two_phase_bytes | Above this estimate of the bytes a one-phase read across nodes would ship, the read uses the two-phase mode instead (ids and distances first, rows second). Formerly [vector] distributed_two_phase_bytes, a name that keeps working (see Renamed keys) |
Every byte-size field accepts the same grammar as the rest of the config file (a plain
integer, or "4GB"/"512MB"/etc; see Config File: File format).
An unrecognized key under [vector], or a value outside the ranges above, is a hard
parse or validation error naming the exact field, never a silently clamped or ignored
setting.
[vector]
memory_budget_total = "8GB"
index_memory_budget = "512MB"
over_fetch = 4
simd_level = "avx2"
ALTER SYSTEM SET vector.<key>
Every key above (and every per-index search knob) is also a live runtime setting, changeable without a restart:
ALTER SYSTEM SET vector.over_fetch = 4;
ALTER SYSTEM RESET vector.over_fetch;
SELECT pg_reload_conf();
ALTER SYSTEM in this release accepts vector.<key> only (superuser required); every
other section is a named follow-up. Three keys take a thousandths integer here instead
of the config file's plain fraction, since a runtime setting cell is always an integer:
| Config-file key (fraction) | ALTER SYSTEM / per-index key (integer, 0-1000) |
|---|---|
selectivity_boundary | vector.selectivity_boundary_permille |
threshold_slack | vector.threshold_slack_permille |
compact_tombstone_ratio | vector.compact_tombstone_ratio_permille |
Every other key keeps its config-file name and form, vector.memory_budget_total
included (it accepts the same "25%" or "8GB" text ALTER SYSTEM SET gives it as a
config file does).
Precedence, highest first: a session parameter (SET anode.<key> = ..., statement
scope) > an index option (ALTER INDEX ... SET) > a catalog override (ALTER SYSTEM SET) > this config file > the built-in default. SHOW vector.<key> and the
scram_vector_settings system view (see Observability)
report both the value in force and which of those five layers set it.
Node-scoped keys: memory_budget_total, maintenance_workers, build_workers,
rerank_prefetch, simd_level, and sample_min_interval_ms are stored per node in a
cluster rather than cluster-wide. ALTER SYSTEM SET vector.<key> = <value> ON NODE '<id>' targets one node; the plain form (no ON NODE) targets the issuing node for
these six keys, and every node for the rest. Every other section of ALTER SYSTEM is
refused in this release, naming the reason.
Docker and cluster config files
The [vector] table follows the same "absent = every field at its default" rule in
every config file ScramDB ships or reads: docker/default-config.toml,
docker/cluster-config.template.toml (rendered by the container entrypoint),
cluster-config.example.toml, local-config.toml, and the Kubernetes ConfigMap that
feeds a Helm-deployed node's config file. Set only the keys you actually want to
override; the rest inherit their built-in defaults exactly as an omitted [vector]
table would. See Clustering and
Kubernetes for where each of those files lives in a real
deployment.
Related pages
- Statements: Vector Indexes -
CREATE INDEX ... WITH (...),ALTER INDEX,REINDEX. - Vector Search - the query patterns these settings tune.
- Tuning: Vector Indexes - what each knob trades off, and in what order to reach for them.
- Observability - the metrics and views that show the value and source of every setting above.
- Config File - every other
[section]this file format accepts.