Skip to main content

Vector index tuning

By the end of this page you will know what each anode/manode search knob trades off, in what order to reach for them, how the filter-tier boundaries pick a plan, and when to run REINDEX instead of turning a knob. This page is about mechanism and direction, not about a target number: how much recall a given knob setting buys, and at what latency and memory cost, is being measured now as those runs complete. Nothing on this page states a recall or latency number that isn't there; treat any number you see elsewhere for these knobs as unverified until it is published with the run it came from.

For where to set each knob (CREATE INDEX ... WITH (...), ALTER INDEX ... SET, a session parameter, or [vector]/ALTER SYSTEM), see Statements and Configuration: Vector Indexes. This page assumes you already have an index and want to change how it searches.

The knobs, in the order to reach for them​

Every level below only widens what the previous one already searched; going straight to the last lever without trying the earlier ones spends latency and memory on a wider search when a cheaper change might have closed the same recall gap.

  1. over_fetch (default 2). The cheapest lever: it asks the index for more candidates before the exact re-rank ever runs, at a cost that scales with the candidate count, not with a wider index walk. Raise this first when a search is close to the recall you want but not quite there.
  2. probe_permille, near_permille (defaults 100, 250). How much of the coarse structure a search actually visits, and how much of that visit is weighted toward the cells nearest the query. Raising either widens the search's reach into the corpus before the graph walk narrows it down.
  3. walk_budget, beam, rank_head (defaults 900, 64, 256). How exhaustively the graph walk explores from the cells the step above selected. These cost more per unit of extra recall than the coarse-search knobs above, since they run the walk itself for longer rather than just giving it a wider starting point.
  4. rerank_extra (default 8). Extra candidates carried into the exact re-rank beyond what k and over_fetch alone would ask for, a final safety margin once the walk itself is already wide.
  5. iterative_scan (default strict_order) and max_scan_tuples (default 20000). What happens to a search that still comes back underfilled after the knobs above: by default the index call itself is widened up to the max_scan_tuples cap rather than accept fewer than k rows. off gives pgvector's behaviour, a possibly short answer with no widening. See Vector Search: Iterative scan modes for the ordering guarantee each mode gives.

Every one of these is a session parameter (SET anode.<key> = ..., tried against one connection with no other effect), a per-index option (ALTER INDEX ... SET (<key> = ...), live immediately, no rebuild), and a [vector] config key / ALTER SYSTEM SET vector.<key> server-wide default, in that precedence order. Try a change as a session parameter first; promote it to ALTER INDEX once you know it helps your workload, and to the server-wide default only once it helps every index that inherits it.

The tier boundaries​

A search with a same-table WHERE predicate picks one of three plans by estimated selectivity (exact_tier_ratio, selectivity_boundary; see Vector Search: Filtering and the three tiers for what each tier does). Both boundaries are settings, not fixed constants, and both ship at a measured prior rather than a proven-universal optimum:

  • exact_tier_ratio (default 4): raise it to send more borderline-selective queries through the exact scan instead of the index, when your table is small enough or your predicates selective enough that the exact path is already fast.
  • selectivity_boundary (default 0.55): raise it to keep more queries on the allow-list traversal tier (one index call, no widening) rather than falling through to the post-filter tier's widening loop.

EXPLAIN on any indexed query names the tier chosen and the selectivity estimate that chose it, so you can see which side of a boundary your real query traffic falls on before you move it.

REINDEX after drift​

An arena's coarse structure is fit once, at build time (and again at each regrow); it does not automatically refit as the data underneath it changes shape. Two signals say it is time to rebuild rather than turn a search knob:

  • Tombstone accumulation. Deleted and updated rows leave tombstones in the arena until compaction runs; scram_vector_indexes.tombstones and the scramdb_vector_index_tombstones metric show the count (see Observability).
  • Distribution shift. The corpus's shape has moved on from what the coarse structure was fit against; scram_vector_indexes.estimator_agreement (the share of candidates whose estimated order matched their exact order in recent searches) trending down, or scram_vector_indexes.recall_at_10 falling below the floor your own recall runs established, is the signal, not a fixed schedule.
CREATE TABLE articles (id BIGINT PRIMARY KEY, embedding VECTOR(3));
INSERT INTO articles VALUES (1, '[0.1,0.2,0.3]'), (2, '[0.9,0.1,0.0]');
CREATE INDEX articles_embedding_idx ON articles USING anode (embedding);

REINDEX INDEX articles_embedding_idx;

DROP TABLE articles;

REINDEX builds a fresh, right-sized arena from the table's current rows and swaps it in atomically: queries in flight keep serving from the old arena until the swap, so there is no window where the index is missing. It also re-sizes the arena promise from the live row count (with arena_slack_percent headroom) rather than the promise the index was originally created with, which is the other reason to reach for it after a table has grown or shrunk substantially since CREATE INDEX.

Memory​

Every vector index on a node reserves its memory ceiling ([vector] index_memory_budget, or the index's own WITH (memory_budget = ...)) from one node-wide total, vector.memory_budget_total, before it writes a file: at CREATE INDEX, at REINDEX, at a regrow, and when a node opens its indexes at start. The reservations together never pass the total. One that does not fit is refused with the arithmetic, for example index needs 1024 MiB, 512 MiB of the vector budget remain, and counted in scramdb_vector_memory_budget_refusals_total; raise vector.memory_budget_total (it applies at once) or give the index a smaller memory_budget. An index that does not fit when its node starts opens in the invalid state (scram_vector_indexes.state) until a REINDEX with room rebuilds it. A write never fails for want of room: past its ceiling a row waits in the index's tail, served by an exact scan merged into every search, until a regrow gives it room.

The decision reads the reservations alone. What the indexes actually use is measured and reported beside them: scram_vector_memory shows the node's total, reserved and in_use (the larger of what the indexes report for themselves and what the allocator measured for vector index and index build memory), the same figures are on /metrics as scramdb_vector_memory_*, and scram_vector_indexes shows each index's own ceiling and use. A measured figure above the reservations never refuses work by itself.

What is not measured yet​

Every recall and latency number for these knobs (how much recall probe_permille = 200 buys over the default, what it costs in QPS, where the tier boundaries actually sit for a given selectivity distribution) is being measured now, not assumed, and will be published here once verified: a number from a real run on a real box, never a guess. This page will gain a worked "raise probe_permille from 100 to 200, recall goes from x to y, latency from p to q" example once that run exists; until then, treat every knob above as a documented lever with a known direction, not a promised number.