Skip to main content

ANODE crushes Qdrant by 20x and Elastic's vector index by 71x

Qdrant and Elastic have spent this season publishing benchmarks at each other. Elastic says Elasticsearch does 32.4 queries a second and Qdrant does 4.5. Qdrant says Qdrant does 111.9. They are 25 times apart on the same system, on the same corpus, and they are fighting about it at two digits.

We ran their benchmark. Same corpus, same queries, same node shape, same load model, ground truth by brute force.

2,318.5 queries a second.

The scoreboard they are fighting over​

The corpus is wiki_dpr_e5: 21,015,300 vectors, 768 dimensions, cosine. The load model is four clients hammering back to back for two minutes per point, recall@100 scored against the exact top 100 of the whole corpus. Three shards, one replica each, a client in the same subnet. Every number below was produced that way.

Their published points, each measured by the vendor that owns the engine. The two Qdrant rows are Qdrant's own; the Elasticsearch row is Elastic's own figure, which Qdrant quoted rather than re-ran. That same Elastic post puts Qdrant at 4.5 queries a second:

SystemNodesQPSrecall@100meanp99
Qdrant 1.18.23 x m6g.xlarge, 4 vCPU111.90.9635.7 ms83.1 ms
Qdrant 1.18.23 x m6g.large, 2 vCPU67.20.9659.5 ms86.0 ms
Elasticsearch DiskBBQ3 x n4-standard-8, 7 vCPU, RF 232.40.96122.6 ms184.3 ms

And ANODE, on the smaller of those two machines, the 4 vCPU one:

ANODEQPSrecall@100meanp99
fast point2,318.50.87061.7 ms2.1 ms
recall 0.90 point1,383.30.89992.9 ms3.4 ms

And it was answering the whole time it was being built. There is no index build step here: the index is complete when the last insert returns, and it takes writes while it serves.

This is their entire product. Qdrant is a vector database and nothing else. Elastic is retrofitting vector search into an engine it built for text, and shipping a new disk format to do it. Between them they published two numbers for the same system that are 25 times apart, and the winner of that argument still answers 111.9 queries a second.

That is 20.7x Qdrant's best published throughput, and 71.6x Elastic's, at 0.87. At recall 0.90 it is 12.4x Qdrant and 42.7x Elastic. Push all the way up to their 0.96 and ANODE still runs close to four times Qdrant's throughput with a p99 nearly seven times lower. There is no operating point in this sweep where the answer is close.

Drag the target and watch what a query costs at each recall:

One query, end to end​

A vector search costs whatever one query costs, and this is where the gap stops being a benchmark artifact.

Same picture from the other side. Their fast configuration answers one query in 35.7 ms on average and 83.1 ms at the ninety-ninth percentile. Elastic's answers in 122.6 ms and 184.3 ms. ANODE answers in 1.7 ms, and its p99 sits within half a millisecond of its mean, because nothing in the read path waits on anything else: no fork to the shards, no join, no straggler.

Why it is not close​

Everybody's vector index does the same two things: narrow the corpus down, then look carefully at what survives. The difference is in who pays for the narrowing.

No fan-out. A sharded index forks a query to every shard and joins the answers, so its latency is a maximum over shards and the slowest one sets the price. ANODE does not fork. A query is answered end to end against one global ordering of the space, and it stops the moment its budget proves that nothing left can beat what it already holds. No fork, no join, no straggler, no cross-shard barrier.

Two movements, one bound. There are two ways to reach a neighbour and they have opposite cost curves: scanning a neighbourhood is cheap per candidate and grows with the corpus, while following a graph costs a fixed number of hops however large the corpus gets. Neither wins alone, so a query uses both, under a single shared bound, so neither half can be widened at the other's expense.

Precision where it decides coverage. Candidates are narrowed in cheap passes, and only the survivors are ever scored exactly. Precision is spent where it changes the answer and saved everywhere it would only cost bandwidth.

The ceiling is an input, not an outcome​

ANODE never uses more memory than you gave it.

The ceiling is a number you set, and the index is built to live inside it. It is not a cache hint the process may drift past, and it does not turn into a surprise at three in the morning. Ask for something that cannot fit, and it tells you at startup rather than an hour into production.

Most vector systems take a cache size instead. The resident set is then a consequence rather than a promise, and the only way to enforce a limit from outside is to kill the process.

The ceiling is configurable, and every number on this page was measured with one in place, on nodes whose memory was a fraction of the corpus. Nobody had to tune a cache to make that work.

Ingestion: 21,005,300 rows in 1,590.9 seconds, 13,203 rows a second, from one client over the network. The index is complete when the last insert returns. There is no background indexing phase to wait out, and neither published post states an ingest time at all.

We do not have an index build time. There is nothing left to do after the last insert returns, which is why this post has an ingest number to print and theirs do not.

There is no build phase to wait out​

This is the part most vector indexes quietly charge you for. You load, then you build, then you wait, then you query, and when the corpus moves you do it again. A rebuild is a window where the index is stale, or offline, or serving two versions of the truth.

ANODE has no such window. The index is complete when the last insert returns. Twenty-one million vectors went in at 13,203 rows a second, and the first query after the last insert was answered by an index that already contained it. Neither published post states an ingest or build time at all.

Writes are not a separate mode. Inserts, updates and deletes run concurrently against the same structure while queries run, and a write is visible to the very next query. The engine's own contract is that an update is searchable within five seconds, and the test suite fails the build when it is not.

Concurrent, and measured while busy. Four search threads and an update thread against one index, fifteen second windows, with a compactor alternately idle and running twenty-one compactions per window:

windowsearches/ssearch p50search p99.9updates/supdate p50
quiet15,051 to 15,562237 to 245 us2.15 to 2.27 ms21,486 to 23,78236 to 37 us
compacting14,962 to 15,225244 to 245 us2.26 to 2.28 ms21,946 to 22,71437 us

Fifteen thousand searches and twenty-two thousand updates a second, on the same index, at the same time, and a compaction running through it costs nothing you can see above the host's own noise.

That is a property of how mutation is arranged, not a scheduling trick. Writers working on different parts of the index do not contend with each other, and nothing is freed while a reader might still be looking at it, so the read path carries no cleanup traffic at all.

Maintenance runs beside the traffic, not instead of it. Every number on this page was measured on a settled index, so none of them is a maintenance-depressed result. When maintenance is running, its cost is bounded and measured: a 500,000-vector re-filing beside live traffic raises search p50 by about half and the p99 into tens of milliseconds for roughly ten seconds, while throughput holds. Pacing that work is the next lever we are pulling, and it is the one place where this engine still spends latency it does not have to. Recall is not what maintenance moves: recall is the budget dial, and the same index answers at 0.9615 when you widen the rerank and the candidate budget.

Maintenance stays out of the way by construction. Fits, refits and the graph backfill run on the index's own thread at background priority, and a streaming write never waits for one. While the index reorganizes itself, queries keep being answered against a consistent view of it, and recall through that window is the settled rate.

Always available is not a slogan here. There is no phase in which this index is not answering.

Vectors belong in the database, not beside it​

Here is the part that outlives the benchmark. Today you run a database for your rows and a vector database for your embeddings, and a pipeline to keep the two agreeing. That is two systems, two failure modes, two bills, and a join your application has to do by hand. When the pipeline is behind, your search is wrong and nothing tells you.

ScramDB does not have that seam. ANODE is an index inside the engine that already holds the rows, so a vector is another column and a nearest-neighbour search is another index scan: the same transaction, the same row-level security, the same branching and point-in-time recovery, the same backup. You filter by tenant and search by vector in one statement, and the answer is consistent because there is nothing to keep in sync.

The surface is pgvector-shaped on purpose. Code written against pgvector keeps working; what changes underneath it is the index that answers, and the index that answers is the one measured on this page.

What this is, and what it is not​

ANODE is the index behind ScramDB's vector search. The SQL surface, pgvector-compatible, is still landing; the engine underneath it is what ran this benchmark.

The honest borders of the comparison: every rival number here is the vendor's own measurement of its own engine, quoted from its own post and not re-run by us. Nobody outside Elastic can re-run DiskBBQ, and Qdrant did not: they quoted it too. Elastic's were produced on different hardware in a different cloud with two replicas, and we quote them as published. Our points come from our own rig on the nodes Qdrant published on, and the two headline points above are cheaper in recall than their 0.96, which is why each one is printed with the recall it was measured at.

Ground truth is the exact top 100 by cosine over all 21 million vectors, computed once by brute force and hashed. Recall reproduces to four decimals between runs. Throughput does not: it moves with the host, and the sweep above is one run of it.

When you can have it​

ANODE ships inside ScramDB in the next iteration, as the index behind pgvector-shaped vector search: same database, same transactions, same row-level security, same backups, no second system to keep in sync.

We did not tune it for this benchmark. We pointed it at their corpus, on their node shape, with their query set, and it answered.