Vector index tuning
By the end of this page you will know what each anode/manode search knob trades
off, in what order to reach for them, how the filter-tier boundaries pick a plan, and
when to run REINDEX instead of turning a knob. This page is about mechanism and
direction, not about a target number: how much recall a given knob setting buys, and at
what latency and memory cost, is being measured now as those runs complete. Nothing on
this page states a recall or latency number that isn't there; treat any number you see
elsewhere for these knobs as unverified until it is published with the run it came
from.
For where to set each knob (CREATE INDEX ... WITH (...), ALTER INDEX ... SET, a
session parameter, or [vector]/ALTER SYSTEM), see
Statements
and Configuration: Vector Indexes. This page assumes you
already have an index and want to change how it searches.
The knobs, in the order to reach for them
Every level below only widens what the previous one already searched; going straight to the last lever without trying the earlier ones spends latency and memory on a wider search when a cheaper change might have closed the same recall gap.
over_fetch(default2). The cheapest lever: it asks the index for more candidates before the exact re-rank ever runs, at a cost that scales with the candidate count, not with a wider index walk. Raise this first when a search is close to the recall you want but not quite there.probe_permille,near_permille(defaults100,250). How much of the coarse structure a search actually visits, and how much of that visit is weighted toward the cells nearest the query. Raising either widens the search's reach into the corpus before the graph walk narrows it down.walk_budget,beam,rank_head(defaults900,64,256). How exhaustively the graph walk explores from the cells the step above selected. These cost more per unit of extra recall than the coarse-search knobs above, since they run the walk itself for longer rather than just giving it a wider starting point.rerank_extra(default8). Extra candidates carried into the exact re-rank beyond whatkandover_fetchalone would ask for, a final safety margin once the walk itself is already wide.iterative_scan(defaultstrict_order) andmax_scan_tuples(default20000). What happens to a search that still comes back underfilled after the knobs above: by default the index call itself is widened up to themax_scan_tuplescap rather than accept fewer thankrows.offgives pgvector's behaviour, a possibly short answer with no widening. See Vector Search: Iterative scan modes for the ordering guarantee each mode gives.
Every one of these is a session parameter (SET anode.<key> = ..., tried against one
connection with no other effect), a per-index option (ALTER INDEX ... SET (<key> = ...), live immediately, no rebuild), and a [vector] config key /
ALTER SYSTEM SET vector.<key> server-wide default, in that precedence order. Try a
change as a session parameter first; promote it to ALTER INDEX once you know it
helps your workload, and to the server-wide default only once it helps every index that
inherits it.
The tier boundaries
A search with a same-table WHERE predicate picks one of three plans by estimated
selectivity (exact_tier_ratio, selectivity_boundary; see Vector Search: Filtering
and the three tiers
for what each tier does). Both boundaries are settings, not fixed constants, and both
ship at a measured prior rather than a proven-universal optimum:
exact_tier_ratio(default4): raise it to send more borderline-selective queries through the exact scan instead of the index, when your table is small enough or your predicates selective enough that the exact path is already fast.selectivity_boundary(default0.55): raise it to keep more queries on the allow-list traversal tier (one index call, no widening) rather than falling through to the post-filter tier's widening loop.
EXPLAIN on any indexed query names the tier chosen and the selectivity estimate that
chose it, so you can see which side of a boundary your real query traffic falls on
before you move it.
REINDEX after drift
An arena's coarse structure is fit once, at build time (and again at each regrow); it does not automatically refit as the data underneath it changes shape. Two signals say it is time to rebuild rather than turn a search knob:
- Tombstone accumulation. Deleted and updated rows leave tombstones in the arena
until compaction runs;
scram_vector_indexes.tombstonesand thescramdb_vector_index_tombstonesmetric show the count (see Observability). - Distribution shift. The corpus's shape has moved on from what the coarse
structure was fit against;
scram_vector_indexes.estimator_agreement(the share of candidates whose estimated order matched their exact order in recent searches) trending down, orscram_vector_indexes.recall_at_10falling below the floor your own recall runs established, is the signal, not a fixed schedule.
CREATE TABLE articles (id BIGINT PRIMARY KEY, embedding VECTOR(3));
INSERT INTO articles VALUES (1, '[0.1,0.2,0.3]'), (2, '[0.9,0.1,0.0]');
CREATE INDEX articles_embedding_idx ON articles USING anode (embedding);
REINDEX INDEX articles_embedding_idx;
DROP TABLE articles;
REINDEX builds a fresh, right-sized arena from the table's current rows and swaps it
in atomically: queries in flight keep serving from the old arena until the swap, so
there is no window where the index is missing. It also re-sizes the arena promise from
the live row count (with arena_slack_percent headroom) rather than the promise the
index was originally created with, which is the other reason to reach for it after a
table has grown or shrunk substantially since CREATE INDEX.
Memory
Every vector index on a node reserves its memory ceiling ([vector] index_memory_budget,
or the index's own WITH (memory_budget = ...)) from one node-wide total,
vector.memory_budget_total, before it writes a file: at CREATE INDEX, at REINDEX,
at a regrow, and when a node opens its indexes at start. The reservations together never
pass the total. One that does not fit is refused with the arithmetic, for example
index needs 1024 MiB, 512 MiB of the vector budget remain, and counted in
scramdb_vector_memory_budget_refusals_total; raise vector.memory_budget_total
(it applies at once) or give the index a smaller memory_budget. An index that does not
fit when its node starts opens in the invalid state (scram_vector_indexes.state) until
a REINDEX with room rebuilds it. A write never fails
for want of room: past its ceiling a row waits in the index's tail, served by an exact
scan merged into every search, until a regrow gives it room.
The decision reads the reservations alone. What the indexes actually use is measured and
reported beside them: scram_vector_memory shows the node's total, reserved and
in_use (the larger of what the indexes report for themselves and what the allocator
measured for vector index and index build memory), the same figures are on /metrics as
scramdb_vector_memory_*, and scram_vector_indexes shows each index's own ceiling and
use. A measured figure above the reservations never refuses work by itself.
What is not measured yet
Every recall and latency number for these knobs (how much recall probe_permille = 200
buys over the default, what it costs in QPS, where the tier boundaries actually sit for
a given selectivity distribution) is being measured now, not assumed, and will be
published here once verified: a number from a real run on a real box, never a guess.
This page will gain a worked "raise probe_permille from 100 to 200, recall goes from
x to y, latency from p to q" example once that run exists; until then, treat
every knob above as a documented lever with a known direction, not a promised number.