Skip to main content

ANODE vector index

ScramDB stores embeddings as a native column type (VECTOR, HALFVEC, SPARSEVEC, and pgvector's BIT distance operators), compatible with pgvector's SQL surface, and indexes them for approximate nearest-neighbour search with ANODE, an in-memory vector index engine built into the workspace rather than bolted on as a plugin. Query it the same way you query any other column: ORDER BY embedding <-> query LIMIT k with an ordinary WHERE, JOIN, or aggregate beside it, on the same live table your application already writes to. See Vector Search for the query surface and Data Types for the column types.

One arena, not a rebuild treadmill​

An ANODE index lives in one memory arena that grows in place as you write: a fresh CREATE INDEX promises capacity for the table's current row count, and later writes that exceed the promise trigger a background regrow to a larger arena rather than a blocking rebuild. Deletes and updates leave tombstones that a background compaction reclaims. Every write is searchable by the very next query; there is no separate "reindex to see recent writes" step.

Two access methods, one engine​

CREATE INDEX ... USING anode builds a single-node index; USING manode builds a sharded, distributed one for a distributed table, one arena per bucket group. Both run the same ANODE engine underneath, so a workload that starts single-node and later moves to a cluster keeps the same index type, options, and query plans. Existing pgvector DDL needs only the method name changed (hnsw/ivfflat become anode/manode); every opclass keeps its pgvector name, since it names the distance, not the engine.

Exact where it says exact, approximate where it says approximate​

Every ANN search re-ranks its candidates against the real stored row before returning them: the index can miss a true neighbour, but it never reports a wrong distance for a row it does return. A query the index cannot yet serve (still building, mid-regrow, or turned off for the session) falls back to an exact scan automatically, with the same answer shape either way. EXPLAIN names which path a query took and, when the index was skipped, why. See Limitations: Vector search for what each metric's index promises and what it does not.

Six metrics, one engine​

L2, cosine, inner product, and L1 distance all run through the same approximate search path; Hamming and Jaccard distance over BIT columns use a separate search path built for bit vectors. Every metric's search operators are compiled for both ScramDB's interpreted and JIT-compiled execution tiers, so a query plan behaves identically whichever tier answers it. See Vector Search: Metric notes per opclass for what each one trades off.

Bounded by construction​

Every vector index arena on a node draws from one shared memory ceiling (vector.memory_budget_total), never the whole machine's memory by surprise. A write that would exceed an index's own budget is never lost: the row enters the index's tail state, served by an exact scan merged into every search, until a regrow gives it room. scram_vector_indexes and the scramdb_vector_* Prometheus family report memory, state, recall, and staleness per index in real time; see Observability.