ANODE vector index
ScramDB stores embeddings as a native column type (VECTOR, HALFVEC, SPARSEVEC,
and pgvector's BIT distance operators), compatible with pgvector's SQL surface, and
indexes them for approximate nearest-neighbour search with ANODE, an in-memory
vector index engine built into the workspace rather than bolted on as a plugin. Query
it the same way you query any other column: ORDER BY embedding <-> query LIMIT k
with an ordinary WHERE, JOIN, or aggregate beside it, on the same live table your
application already writes to. See Vector Search
for the query surface and Data Types for
the column types.
One arena, not a rebuild treadmill
An ANODE index lives in one memory arena that grows in place as you write: a fresh
CREATE INDEX promises capacity for the table's current row count, and later writes
that exceed the promise trigger a background regrow to a larger arena rather than a
blocking rebuild. Deletes and updates leave tombstones that a background compaction
reclaims. Every write is searchable by the very next query; there is no separate
"reindex to see recent writes" step.
Two access methods, one engine
CREATE INDEX ... USING anode builds a single-node index; USING manode builds a
sharded, distributed one for a distributed table, one arena per bucket group. Both run
the same ANODE engine underneath, so a workload that starts single-node and later moves
to a cluster keeps the same index type, options, and query plans. Existing pgvector DDL
needs only the method name changed (hnsw/ivfflat become anode/manode); every
opclass keeps its pgvector name, since it names the distance, not the engine.
Exact where it says exact, approximate where it says approximate
Every ANN search re-ranks its candidates against the real stored row before returning
them: the index can miss a true neighbour, but it never reports a wrong distance for a
row it does return. A query the index cannot yet serve (still building, mid-regrow,
or turned off for the session) falls back to an exact scan automatically, with the same
answer shape either way. EXPLAIN names which path a query took and, when the index
was skipped, why. See Limitations: Vector search
for what each metric's index promises and what it does not.
Six metrics, one engine
L2, cosine, inner product, and L1 distance all run through the same approximate
search path; Hamming and Jaccard distance over BIT columns use a separate search
path built for bit vectors. Every metric's search operators are compiled for both
ScramDB's interpreted and JIT-compiled execution tiers, so a query plan behaves
identically whichever tier answers it. See Vector Search: Metric notes per
opclass for what each one
trades off.
Bounded by construction
Every vector index arena on a node draws from one shared memory ceiling
(vector.memory_budget_total), never the whole machine's memory by surprise. A write
that would exceed an index's own budget is never lost: the row enters the index's
tail state, served by an exact scan merged into every search, until a regrow gives it
room. scram_vector_indexes and the scramdb_vector_* Prometheus family report memory,
state, recall, and staleness per index in real time; see
Observability.