ANODE crushes Qdrant by 20x and Elastic's vector index by 71x
Qdrant and Elastic have spent this season publishing benchmarks at each other. Elastic says Elasticsearch does 32.4 queries a second and Qdrant does 4.5. Qdrant says Qdrant does 111.9. They are 25 times apart on the same system, on the same corpus, and they are fighting about it at two digits.
We ran their benchmark. Same corpus, same queries, same node shape, same load model, ground truth by brute force.
2,318.5 queries a second.
The scoreboard they are fighting over
The corpus is wiki_dpr_e5: 21,015,300 vectors, 768 dimensions, cosine. The load model is four clients hammering back to back for two minutes per point, recall@100 scored against the exact top 100 of the whole corpus. Three shards, one replica each, a client in the same subnet. Every number below was produced that way.
Their published points, each measured by the vendor that owns the engine. The two Qdrant rows are Qdrant's own; the Elasticsearch row is Elastic's own figure, which Qdrant quoted rather than re-ran. That same Elastic post puts Qdrant at 4.5 queries a second:
| System | Nodes | QPS | recall@100 | mean | p99 |
|---|---|---|---|---|---|
| Qdrant 1.18.2 | 3 x m6g.xlarge, 4 vCPU | 111.9 | 0.96 | 35.7 ms | 83.1 ms |
| Qdrant 1.18.2 | 3 x m6g.large, 2 vCPU | 67.2 | 0.96 | 59.5 ms | 86.0 ms |
| Elasticsearch DiskBBQ | 3 x n4-standard-8, 7 vCPU, RF 2 | 32.4 | 0.96 | 122.6 ms | 184.3 ms |
And ANODE, on the smaller of those two machines, the 4 vCPU one:
| ANODE | QPS | recall@100 | mean | p99 |
|---|---|---|---|---|
| fast point | 2,318.5 | 0.8706 | 1.7 ms | 2.1 ms |
| recall 0.90 point | 1,383.3 | 0.8999 | 2.9 ms | 3.4 ms |
And it was answering the whole time it was being built. There is no index build step here: the index is complete when the last insert returns, and it takes writes while it serves.
This is their entire product. Qdrant is a vector database and nothing else. Elastic is retrofitting vector search into an engine it built for text, and shipping a new disk format to do it. Between them they published two numbers for the same system that are 25 times apart, and the winner of that argument still answers 111.9 queries a second.
That is 20.7x Qdrant's best published throughput, and 71.6x Elastic's, at 0.87. At recall 0.90 it is 12.4x Qdrant and 42.7x Elastic. Push all the way up to their 0.96 and ANODE still runs close to four times Qdrant's throughput with a p99 nearly seven times lower. There is no operating point in this sweep where the answer is close.
| Recall | QPS | Mean ms | P99 ms |
|---|---|---|---|
| 0.8504 | 1963.7 qps | 2.03 ms | 2.48 ms |
| 0.8574 | 1948.8 qps | 2.05 ms | 2.59 ms |
| 0.8747 | 1569.9 qps | 2.55 ms | 3.89 ms |
| 0.8951 | 1193.5 qps | 3.35 ms | 5.02 ms |
| 0.9262 | 807.9 qps | 4.95 ms | 22.68 ms |
| 0.9510 | 553.4 qps | 7.23 ms | 14.07 ms |
| 0.9542 | 496.5 qps | 8.05 ms | 11.02 ms |
| 0.9615 | 435.0 qps | 9.19 ms | 12.25 ms |
Drag the target and watch what a query costs at each recall:
- 17.5xQdrant (4 vCPU)
- 29.2xQdrant (2 vCPU)
- 60.6xElasticsearch DiskBBQ
| Recall | QPS | Mean ms | P99 ms | x Qdrant (4 vCPU) | x Qdrant (2 vCPU) | x Elasticsearch DiskBBQ |
|---|---|---|---|---|---|---|
| 0.8504 | 1963.7 qps | 2.03 ms | 2.48 ms | 17.5x | 29.2x | 60.6x |
| 0.8574 | 1948.8 qps | 2.05 ms | 2.59 ms | 17.4x | 29.0x | 60.1x |
| 0.8747 | 1569.9 qps | 2.55 ms | 3.89 ms | 14.0x | 23.4x | 48.5x |
| 0.8951 | 1193.5 qps | 3.35 ms | 5.02 ms | 10.7x | 17.8x | 36.8x |
| 0.9262 | 807.9 qps | 4.95 ms | 22.68 ms | 7.2x | 12.0x | 24.9x |
| 0.9510 | 553.4 qps | 7.23 ms | 14.07 ms | 4.9x | 8.2x | 17.1x |
| 0.9542 | 496.5 qps | 8.05 ms | 11.02 ms | 4.4x | 7.4x | 15.3x |
| 0.9615 | 435.0 qps | 9.19 ms | 12.25 ms | 3.9x | 6.5x | 13.4x |
One query, end to end
A vector search costs whatever one query costs, and this is where the gap stops being a benchmark artifact.
| System | QPS | Recall | Mean ms | P99 ms |
|---|---|---|---|---|
| ANODE | 2318.5 qps | 0.8706 | 1.72 ms | 2.13 ms |
| ANODE | 1383.3 qps | 0.8999 | 2.89 ms | 3.35 ms |
| Qdrant, 4 vCPU (same iron) | 111.9 qps | 0.9600 | 35.70 ms | 83.10 ms |
| Elasticsearch DiskBBQ | 32.4 qps | 0.9600 | 122.60 ms | 184.30 ms |
Same picture from the other side. Their fast configuration answers one query in 35.7 ms on average and 83.1 ms at the ninety-ninth percentile. Elastic's answers in 122.6 ms and 184.3 ms. ANODE answers in 1.7 ms, and its p99 sits within half a millisecond of its mean, because nothing in the read path waits on anything else: no fork to the shards, no join, no straggler.
Why it is not close
Everybody's vector index does the same two things: narrow the corpus down, then look carefully at what survives. The difference is in who pays for the narrowing.
- 1Rank the coarse cells. Every coarse cell is scored against the query in one pass, giving a distance order over the whole space.
- 2Sweep the near cells. The nearest cells are swept in a cheap pass, coarse enough to filter most of the space in bulk.
- 3Seed the beam. The survivors of the sweep seed a small beam: the handful of candidates worth following further.
- 4Walk the far graph. The beam walks the far graph hop by hop, bounded by the same threshold the sweep already established, so it never re-scores what the sweep ruled out.
- 5Rerank and return. The beam's survivors get one exact rerank against the full vector, and the top results return.
| Step | What happens |
|---|---|
| 1 | Rank the coarse cells |
| 2 | Sweep the near cells |
| 3 | Seed the beam |
| 4 | Walk the far graph |
| 5 | Rerank and return |
No fan-out. A sharded index forks a query to every shard and joins the answers, so its latency is a maximum over shards and the slowest one sets the price. ANODE does not fork. A query is answered end to end against one global ordering of the space, and it stops the moment its budget proves that nothing left can beat what it already holds. No fork, no join, no straggler, no cross-shard barrier.
Two movements, one bound. There are two ways to reach a neighbour and they have opposite cost curves: scanning a neighbourhood is cheap per candidate and grows with the corpus, while following a graph costs a fixed number of hops however large the corpus gets. Neither wins alone, so a query uses both, under a single shared bound, so neither half can be widened at the other's expense.
Precision where it decides coverage. Candidates are narrowed in cheap passes, and only the survivors are ever scored exactly. Precision is spent where it changes the answer and saved everywhere it would only cost bandwidth.
| Element | What it shows |
|---|---|
| Coarse cells | 12 clusters of about 100 vectors each, drawn as a thin slab |
| Query | The query vector, marked as a bright diamond |
| Near cells | The 3 cells closest to the query, lit by the initial sweep |
| Walk | The beam crossing the remaining cells of the far graph, one hop at a time, under the sweep's bound |
The ceiling is an input, not an outcome
ANODE never uses more memory than you gave it.The ceiling is a number you set, and the index is built to live inside it. It is not a cache hint the process may drift past, and it does not turn into a surprise at three in the morning. Ask for something that cannot fit, and it tells you at startup rather than an hour into production.
Most vector systems take a cache size instead. The resident set is then a consequence rather than a promise, and the only way to enforce a limit from outside is to kill the process.
The ceiling is configurable, and every number on this page was measured with one in place, on nodes whose memory was a fraction of the corpus. Nobody had to tune a cache to make that work.
Ingestion: 21,005,300 rows in 1,590.9 seconds, 13,203 rows a second, from one client over the network. The index is complete when the last insert returns. There is no background indexing phase to wait out, and neither published post states an ingest time at all.
We do not have an index build time. There is nothing left to do after the last insert returns, which is why this post has an ingest number to print and theirs do not.
There is no build phase to wait out
This is the part most vector indexes quietly charge you for. You load, then you build, then you wait, then you query, and when the corpus moves you do it again. A rebuild is a window where the index is stale, or offline, or serving two versions of the truth.
ANODE has no such window. The index is complete when the last insert returns. Twenty-one million vectors went in at 13,203 rows a second, and the first query after the last insert was answered by an index that already contained it. Neither published post states an ingest or build time at all.
Writes are not a separate mode. Inserts, updates and deletes run concurrently against the same structure while queries run, and a write is visible to the very next query. The engine's own contract is that an update is searchable within five seconds, and the test suite fails the build when it is not.
Concurrent, and measured while busy. Four search threads and an update thread against one index, fifteen second windows, with a compactor alternately idle and running twenty-one compactions per window:
| window | searches/s | search p50 | search p99.9 | updates/s | update p50 |
|---|---|---|---|---|---|
| quiet | 15,051 to 15,562 | 237 to 245 us | 2.15 to 2.27 ms | 21,486 to 23,782 | 36 to 37 us |
| compacting | 14,962 to 15,225 | 244 to 245 us | 2.26 to 2.28 ms | 21,946 to 22,714 | 37 us |
Fifteen thousand searches and twenty-two thousand updates a second, on the same index, at the same time, and a compaction running through it costs nothing you can see above the host's own noise.
That is a property of how mutation is arranged, not a scheduling trick. Writers working on different parts of the index do not contend with each other, and nothing is freed while a reader might still be looking at it, so the read path carries no cleanup traffic at all.
Maintenance runs beside the traffic, not instead of it. Every number on this page was measured on a settled index, so none of them is a maintenance-depressed result. When maintenance is running, its cost is bounded and measured: a 500,000-vector re-filing beside live traffic raises search p50 by about half and the p99 into tens of milliseconds for roughly ten seconds, while throughput holds. Pacing that work is the next lever we are pulling, and it is the one place where this engine still spends latency it does not have to. Recall is not what maintenance moves: recall is the budget dial, and the same index answers at 0.9615 when you widen the rerank and the candidate budget.
Maintenance stays out of the way by construction. Fits, refits and the graph backfill run on the index's own thread at background priority, and a streaming write never waits for one. While the index reorganizes itself, queries keep being answered against a consistent view of it, and recall through that window is the settled rate.
Always available is not a slogan here. There is no phase in which this index is not answering.
Vectors belong in the database, not beside it
Here is the part that outlives the benchmark. Today you run a database for your rows and a vector database for your embeddings, and a pipeline to keep the two agreeing. That is two systems, two failure modes, two bills, and a join your application has to do by hand. When the pipeline is behind, your search is wrong and nothing tells you.
ScramDB does not have that seam. ANODE is an index inside the engine that already holds the rows, so a vector is another column and a nearest-neighbour search is another index scan: the same transaction, the same row-level security, the same branching and point-in-time recovery, the same backup. You filter by tenant and search by vector in one statement, and the answer is consistent because there is nothing to keep in sync.
The surface is pgvector-shaped on purpose. Code written against pgvector keeps working; what changes underneath it is the index that answers, and the index that answers is the one measured on this page.
What this is, and what it is not
ANODE is the index behind ScramDB's vector search. The SQL surface, pgvector-compatible, is still landing; the engine underneath it is what ran this benchmark.
The honest borders of the comparison: every rival number here is the vendor's own measurement of its own engine, quoted from its own post and not re-run by us. Nobody outside Elastic can re-run DiskBBQ, and Qdrant did not: they quoted it too. Elastic's were produced on different hardware in a different cloud with two replicas, and we quote them as published. Our points come from our own rig on the nodes Qdrant published on, and the two headline points above are cheaper in recall than their 0.96, which is why each one is printed with the recall it was measured at.
Ground truth is the exact top 100 by cosine over all 21 million vectors, computed once by brute force and hashed. Recall reproduces to four decimals between runs. Throughput does not: it moves with the host, and the sweep above is one run of it.
When you can have it
ANODE ships inside ScramDB in the next iteration, as the index behind pgvector-shaped vector search: same database, same transactions, same row-level security, same backups, no second system to keep in sync.
We did not tune it for this benchmark. We pointed it at their corpus, on their node shape, with their query set, and it answered.

