← All releases
Release benchmarks

Antfly 0.2

Relevant information doesn't always look like a nearest neighbor. Antfly combines vector, full-text, and graph indexes with native document processing and inference, so you can choose and combine retrieval methods for your data.

Antfly 0.2 · Benchmark overview

Retrieval takes more than vector search

Six metrics spanning search, relationships, document processing, and inference, all supported by Antfly.

Antfly

Antfly measured benchmark axes
Vector
3.53 ms
Full-text search
5.55 ms
Retrieval quality
49.30%
Filtered graph retrieval
0.530 ms
Text generation
50.00 tok/s
PDF retrieval
44.33%
Details

Eligible functionality is bundled with the self-hosted product version and edition identified in each benchmark section. PostgreSQL + pgvector is evaluated as one combined product, our sole separately installed extension exception.

A “None” reading identifies functionality unavailable under these criteria and centers that axis. Each benchmark section’s Details documents its workload, input preparation, runtime, and client transport.

The PDF axis evaluates native text extraction with a common external retrieval evaluator. Elasticsearch receives externally page-separated PDF bytes; Antfly receives whole PDFs. The full-text axis measures ranked lexical retrieval; Chroma’s self-hosted text API provides document-content filtering.

Matched ground-truth text defines 100% on the PDF axis. Every other axis uses the best listed measurement as its outer ring, with lower-is-better measurements inverted.

01

Retrieval quality

STaRK is a set of three benchmarks, each measuring a different workload. Prime measures document retrieval for specific biomedical questions, Amazon focuses on product recommendations, and MAG retrieves academic papers from a pool of 700,244 paper candidates. Quality is measured by Recall@20: the fraction of expected answers returned in the first 20 results, averaged across queries. Retrieval method matters: Antfly’s hybrid search over relation-expanded text reached 49.30% on Prime, versus 36.00% for dense retrieval. On MAG, keyword search over relation-expanded text reached 72.18%, versus 48.36% for dense retrieval.

Benchmark definition

STaRK-Prime · hybrid over relation-expanded text

Recall@20 · same 2,801 queries · higher is better

Recall@20 · same 2,801 queries · higher is better. Antfly: Antfly, 49.30%. Neo4j: Neo4j, 48.40%. Weaviate: Weaviate, 48.19%. Elasticsearch: Elasticsearch, 43.91%. Postgres + pgvector: Postgres + pgvector, 30.43%. Milvus: Milvus, 46.65%

  • Antfly
  • Other systems

STaRK-Amazon · default-text retrieval

Recall@20 · 1,642 queries · one Antfly run; published baselines are external references

Recall@20 · 1,642 queries · one Antfly run; published baselines are external references. Antfly: Antfly · hybrid, 60.75%; Antfly · keyword, 55.85%. Published reference: BM25, 53.77%; ada-002, 53.29%

  • Antfly (all-term OR)
  • Published baselines

STaRK-MAG · Antfly methods and published baselines

Recall@20 · 2,665 queries · published baselines are external references

Recall@20 · 2,665 queries · published baselines are external references. Antfly: Antfly · keyword, 72.18%; Antfly · hybrid, relation-expanded, 70.38%; Antfly · dense, 48.36%. Published reference: Multi-Ada, 50.80%; Voyage L2, 50.49%

  • Antfly (by method)
  • Published baselines
Details

Chroma, Elasticsearch, Milvus, Neo4j, Postgres + pgvector, Quickwit, and Weaviate use the complete 2,801-query STaRK-Prime test split. The radar uses each system’s best measured Recall@20. The Prime chart shows hybrid retrieval over relation-expanded text; Quickwit’s native BM25 results appear in the radar and mode matrix.

Prime Antfly values are means of two runs; each database peer value comes from one run. Published paper baselines are shown separately from the database comparison.

Chroma contributes its dense-retrieval result to the radar and mode matrix. Its self-hosted text API provides document-content filtering; ranked lexical retrieval is a Cloud feature.

The mode matrix lists every measured Prime retrieval configuration.

Relation-expanded text adds relationship information to each document before indexing. Queries retrieve from this flattened representation.

Amazon uses default-text queries over 957,192 product candidates. Antfly keyword retrieval uses all-term OR and hybrid retrieval uses built-in RRF, measured in one rc2 run (4be5a0e9). BM25 and ada-002 values come from the published Table 6 baselines.

Tested versions

Antfly
Prime: v0.2.1-rc1 · MAG and Amazon: v0.2.1-rc2
Published reference
Original-paper full-test baselines
Chroma
1.0.0
Elasticsearch
9.4.1
Neo4j
5.26.29
Postgres + pgvector
0.8.5 on PostgreSQL 16.14
Weaviate
1.38.2
Milvus
3.0.0
Quickwit
0.9.0

STaRK-Prime Recall@20 · every measured engine and mode

STaRK-Prime Recall@20 · every measured engine and mode
DatabaseDense · official vectorsLexical · default textHybrid · default textLexical · relation-expandedHybrid · relation-expanded
Antfly36.00%33.56%40.37%45.40%49.30%
Neo4j34.42%31.64%39.21%43.06%48.40%
Weaviate35.34%31.77%39.24%43.80%48.19%
Elasticsearch35.78%31.58%32.66%43.06%43.91%
Postgres + pgvector35.55%3.07%29.35%3.60%30.43%
Chroma34.94%UnsupportedUnsupportedUnsupportedUnsupported
QuickwitUnsupported30.85%Unsupported42.65%Unsupported
Milvus32.61%31.76%37.51%43.49%46.65%

Dense retrieval versus the best measured combination

Dense retrieval versus the best measured combination
DatasetDense retrievalBest measured combinationDifference
STaRK-Prime36.00%49.30% · fusion + relationship context+13.30 percentage points
STaRK-MAG48.36%72.18% · lexical + relationship context+23.82 percentage points

Database results used by the radar

Database results used by the radar
DatabaseBest measured methodSTaRK-Prime Recall@20
AntflyHybrid · relation-expanded text49.30%
Neo4jHybrid · relation-expanded text48.40%
WeaviateHybrid · relation-expanded text48.19%
ElasticsearchHybrid · relation-expanded text43.91%
Postgres + pgvectorDense35.55%
ChromaDense34.94%
QuickwitBM25 · relation-expanded text42.65%
MilvusHybrid · relation-expanded text46.65%

Other Antfly retrieval methods

Other Antfly retrieval methods
Antfly methodSTaRK-Prime Recall@20STaRK-MAG Recall@20
Hybrid · relation-expanded text49.30%70.38%
Keyword · relation-expanded text45.40%72.18%
RRF default40.37%52.76%
Dense36.00%48.36%
Lexical33.56%46.11%

Milvus STaRK-Prime · all native modes and controls

Mode
Upstream exact-vector control
hit@1
0.126383
hit@5
0.314888
recall@20
0.360008
mrr
0.209928
Completed queries
2801
Errors
0
Mode
Upstream BM25 control
hit@1
0.127454
hit@5
0.279186
recall@20
0.312506
mrr
0.195229
Completed queries
2801
Errors
0
Mode
BM25 · default text
hit@1
0.127811
hit@5
0.273474
recall@20
0.317602
mrr
0.194232
Completed queries
2801
Errors
0
Mode
Dense · official vectors
hit@1
0.123170
hit@5
0.300607
recall@20
0.326095
mrr
0.201137
Completed queries
2801
Errors
0
Mode
Hybrid · default text
hit@1
0.128169
hit@5
0.331310
recall@20
0.375114
mrr
0.217760
Completed queries
2801
Errors
0
Mode
BM25 · relation-expanded text
hit@1
0.179579
hit@5
0.373081
recall@20
0.434913
mrr
0.266339
Completed queries
2801
Errors
0
Mode
Hybrid · relation-expanded text
hit@1
0.153874
hit@5
0.369511
recall@20
0.466545
mrr
0.252422
Completed queries
2801
Errors
0

Methodology

  • All three datasets use their complete official synthetic test splits and the official Recall@20 metric.
  • Prime contains 2,801 test queries over 129,375 records; MAG contains 2,665 test queries over 700,244 paper candidates within the larger knowledge graph; Amazon contains 1,642 test queries over 957,192 product candidates.
  • Published comparison bars use the original full-test results.
  • Relation-expanded text uses STaRK’s flattened relation-context serialization, materialized before indexing.
  • Postgres + pgvector is one system in these comparisons: PostgreSQL provides full-text search and pgvector provides vector search.
  • Milvus’s supplemental Prime run uses one fresh run, all 129,375 candidates and 2,801 official queries, with zero errors in all five native modes and both upstream controls. Native RRF uses k=60 and two 100-result windows. TEXT fields with Storage V3 and PyMilvus 3.0.1 retain long documents without truncation, splitting or client-side fusion; the server is Milvus 3.0.0.
  • Quickwit 0.9.0 uses native BM25 over all 129,375 candidates and 2,801 queries in one run. Both BM25 modes and both upstream controls completed with zero errors. Recall@20 is 30.8509% for default text and 42.6478% for relation-expanded text. The mode matrix identifies unsupported dense/hybrid modes; latency and resource use are measurements from a single run.
02

Full-text search

These runs compare full-text search as the corpus grows from 250K to 3M passages. MIRACL English supplies 799 queries with 8,350 human relevance judgments, sampled against its 32.9M-passage source corpus.

250K query p95

Logarithmic bar scale · baseline 1 ms

Antfly: Antfly, 2.06 ms. Elasticsearch: Elasticsearch, 3.44 ms. Quickwit: Quickwit, 22.80 ms. Postgres + pgvector: Postgres + pgvector, 99 ms

  • Antfly
  • Other systems

1M query p95

Logarithmic bar scale · baseline 1 ms

Antfly: Antfly, 5.55 ms. Elasticsearch: Elasticsearch, 4.36 ms. Quickwit: Quickwit, 25.66 ms. Postgres + pgvector: Postgres + pgvector, 336 ms. Neo4j: Neo4j, 5.16 ms. Weaviate: Weaviate, 4.58 ms. Milvus: Milvus, 2.88 ms

  • Antfly
  • Other systems

2M query p95

Logarithmic bar scale · baseline 1 ms

Antfly: Antfly, 9.93 ms. Elasticsearch: Elasticsearch, 5.62 ms. Quickwit: Quickwit, 27.82 ms. Postgres + pgvector: Postgres + pgvector, 650 ms

  • Antfly
  • Other systems

3M query p95

Logarithmic bar scale · baseline 1 ms

Antfly: Antfly, 14.37 ms. Elasticsearch: Elasticsearch, 7.15 ms. Quickwit: Quickwit, 29.29 ms

  • Antfly
  • Other systems
Details

Tested versions

Antfly
v0.2.1
Elasticsearch
9.5.2
Quickwit
0.9.0
Postgres + pgvector
PostgreSQL 16 · built-in full-text
Neo4j
5.26.29
Weaviate
1.39.2
Milvus
3.0.0

Working set and post-stop volume

Engine
Antfly
Passages
250,000
Demand (MB)
360.2 MB [335.9–367.5]
Cache + memory (MB)
643.5 MB [636.0–662.8]
Final volume (MB)
160.5 MB [157.0–165.3]
Runs / storage status
3 / 3
Engine
Antfly
Passages
1,000,000
Demand (MB)
1,554.1 MB [1,549.9–1,998.8]
Cache + memory (MB)
2,385.2 MB [2,316.5–2,677.8]
Final volume (MB)
597.9 MB [594.5–601.2]
Runs / storage status
3 / 3
Engine
Antfly
Passages
2,000,000
Demand (MB)
3,606.5 MB [3,264.7–4,415.8]
Cache + memory (MB)
4,884.7 MB [4,874.2–5,736.5]
Final volume (MB)
1,176.6 MB [1,143.6–1,232.7]
Runs / storage status
3 / 3
Engine
Antfly
Passages
3,000,000
Demand (MB)
6,884.3 MB [6,631.1–7,675.8]
Cache + memory (MB)
9,545.9 MB [9,250.7–10,431.8]
Final volume (MB)
1,706.6 MB [1,701.5–1,716.2]
Runs / storage status
3 / 3
Engine
Elasticsearch
Passages
250,000
Demand (MB)
1,712.7 MB
Cache + memory (MB)
2,305.5 MB
Final volume (MB)
123.3 MB
Runs / storage status
1 / 1
Engine
Postgres + pgvector
Passages
250,000
Demand (MB)
207.6 MB
Cache + memory (MB)
901.3 MB
Final volume (MB)
665.9 MB
Runs / storage status
1 / 1
Engine
Quickwit
Passages
250,000
Demand (MB)
547.4 MB
Cache + memory (MB)
767.7 MB
Final volume (MB)
202.7 MB
Runs / storage status
1 / 1
Engine
Elasticsearch
Passages
1,000,000
Demand (MB)
1,801.2 MB
Cache + memory (MB)
2,886.6 MB
Final volume (MB)
477.4 MB
Runs / storage status
1 / 1
Engine
Quickwit
Passages
1,000,000
Demand (MB)
1,102.4 MB
Cache + memory (MB)
1,725.0 MB
Final volume (MB)
390.6 MB
Runs / storage status
1 / 1
Engine
Elasticsearch
Passages
2,000,000
Demand (MB)
1,881.6 MB
Cache + memory (MB)
5,001.3 MB
Final volume (MB)
947.0 MB
Runs / storage status
1 / 1
Engine
Quickwit
Passages
2,000,000
Demand (MB)
1,827.8 MB
Cache + memory (MB)
2,936.2 MB
Final volume (MB)
773.4 MB
Runs / storage status
1 / 1
Engine
Elasticsearch
Passages
3,000,000
Demand (MB)
1,880.9 MB
Cache + memory (MB)
4,903.8 MB
Final volume (MB)
1,416.2 MB
Runs / storage status
1 / 1
Engine
Quickwit
Passages
3,000,000
Demand (MB)
2,210.9 MB
Cache + memory (MB)
3,818.9 MB
Final volume (MB)
1,029.1 MB
Runs / storage status
1 / 1
Engine
Neo4j
Passages
1000000
Demand (MB)
2164.5 MB [2158.5–2183.6]
Cache + memory (MB)
2788.5 MB [2749.0–2802.1]
Final volume (MB)
2938.3 MB [2938.2–2938.3]
Runs / storage status
3 / 3 runs; storage verified
Engine
Weaviate
Passages
1000000
Demand (MB)
3737.1 MB [3296.5–3808.4]
Cache + memory (MB)
5419.7 MB [5327.5–6015.1]
Final volume (MB)
1765.4 MB [1763.7–1766.2]
Runs / storage status
3 / 3 runs; storage verified
Engine
Milvus
Passages
1000000
Demand (MB)
5538.6 MB [5079.3–5875.0]
Cache + memory (MB)
7722.1 MB [7164.1–8057.5]
Final volume (MB)
2646.9 MB [2493.8–2724.7] (provisional)
Runs / storage status
3 / 3 runs; storage provisional

Query and judged retrieval evidence

Engine
Antfly
Passages
250,000
QPS
905.537 [871.046–914.334]
p95 (ms)
2.06 [2.04–2.13]
nDCG@10
0.572 [0.571–0.574]
Recall@100
0.941 [0.938–0.946]
Judged@100
0.097 [0.097–0.098]
Errors
0
Engine
Antfly
Passages
1,000,000
QPS
435.758 [431.344–437.734]
p95 (ms)
5.55 [5.45–5.69]
nDCG@10
0.525 [0.525–0.531]
Recall@100
0.921 [0.912–0.924]
Judged@100
0.094 [0.094–0.094]
Errors
0
Engine
Antfly
Passages
2,000,000
QPS
264.899 [264.051–268.398]
p95 (ms)
9.93 [9.79–10.22]
nDCG@10
0.454 [0.441–0.493]
Recall@100
0.815 [0.803–0.893]
Judged@100
0.084 [0.083–0.091]
Errors
0
Engine
Antfly
Passages
3,000,000
QPS
193.371 [192.881–193.966]
p95 (ms)
14.37 [14.15–14.67]
nDCG@10
0.471 [0.382–0.476]
Recall@100
0.879 [0.744–0.884]
Judged@100
0.090 [0.075–0.090]
Errors
0
Engine
Elasticsearch
Passages
250,000
QPS
363.665
p95 (ms)
3.44
nDCG@10
0.552
Recall@100
0.930
Judged@100
0.094
Errors
0
Engine
Quickwit
Passages
250,000
QPS
47.024
p95 (ms)
22.80
nDCG@10
0.555
Recall@100
0.939
Judged@100
0.095
Errors
0
Engine
Elasticsearch
Passages
1,000,000
QPS
302.605
p95 (ms)
4.36
nDCG@10
0.493
Recall@100
0.892
Judged@100
0.090
Errors
0
Engine
Quickwit
Passages
1,000,000
QPS
43.014
p95 (ms)
25.66
nDCG@10
0.498
Recall@100
0.903
Judged@100
0.091
Errors
0
Engine
Elasticsearch
Passages
2,000,000
QPS
258.106
p95 (ms)
5.62
nDCG@10
0.453
Recall@100
0.866
Judged@100
0.087
Errors
0
Engine
Quickwit
Passages
2,000,000
QPS
40.320
p95 (ms)
27.82
nDCG@10
0.458
Recall@100
0.877
Judged@100
0.088
Errors
0
Engine
Elasticsearch
Passages
3,000,000
QPS
195.861
p95 (ms)
7.15
nDCG@10
0.425
Recall@100
0.849
Judged@100
0.085
Errors
0
Engine
Quickwit
Passages
3,000,000
QPS
38.805
p95 (ms)
29.29
nDCG@10
0.427
Recall@100
0.856
Judged@100
0.086
Errors
0
Engine
Neo4j
Passages
1000000
QPS
276.93 [272.70–281.52]
p95 (ms)
5.16 [5.05–5.32]
nDCG@10
0.4921 [0.4920–0.4922]
Recall@100
0.8904 [0.8904–0.8904]
Judged@100
0.0899
Errors
0
Engine
Weaviate
Passages
1000000
QPS
375.52 [271.16–411.02]
p95 (ms)
4.58 [3.79–7.85]
nDCG@10
0.4960 [0.4959–0.4962]
Recall@100
0.8993 [0.8993–0.8993]
Judged@100
0.0911
Errors
0
Engine
Milvus
Passages
1000000
QPS
414.83 [400.23–456.10]
p95 (ms)
2.88 [2.70–4.22]
nDCG@10
0.4987 [0.4987–0.4987]
Recall@100
0.9037 [0.9037–0.9037]
Judged@100
0.0913
Errors
0

PostgreSQL controlled MIRACL query results

Passages
250,000
First-trial p95 (ms)
99
Query errors
0
Passages
1,000,000
First-trial p95 (ms)
336
Query errors
0
Passages
2,000,000
First-trial p95 (ms)
650
Query errors
0

Methodology

  • Antfly v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449. Recorded host: Apple M4 Max (linux-arm64).
  • Identical nested MIRACL English fixtures at 250K, 1M, 2M, and 3M passages. Antfly uses three fresh runs per size, ten warmup queries, and all 799 measured queries at concurrency one.
  • Database services run in Linux arm64 containers on the recorded host.
  • Neo4j, Weaviate and Milvus 1M-passage results use three fresh Linux ARM64 runs on Apple M4 Max / Podman 6.1.1, all 799 judged queries, top-100 and zero query errors. p95, QPS and serving memory use run medians; quality and final storage use means. Tables retain ranges. Milvus executes BM25 natively on the server.
03

Vector search

How fast is a semantic query at comparable recall? These VectorDBBench runs compare serial p95 latency on 50K 1536-dimensional vectors and 1M 768-dimensional vectors, with quality and resource measurements in Details. Compared with v0.2.0, measured 1M serial p95 falls from 79.67 to 14.90 ms (81.3% lower), with recall 0.960 → 0.963.

50K serial query p95

Antfly: Antfly, 3.533 ms. Chroma: Chroma, 8.067 ms. Elasticsearch: Elasticsearch, 2.033 ms. Postgres + pgvector: Postgres + pgvector, 4.167 ms. Weaviate: Weaviate, 6.300 ms. Milvus: Milvus, 1.700 ms. Neo4j: Neo4j, 28.533 ms

  • Antfly
  • Other systems

1M serial query p95

Antfly: Antfly, 14.900 ms. Chroma: Chroma, 6.600 ms. Elasticsearch: Elasticsearch, 3.233 ms. Postgres + pgvector: Postgres + pgvector, 10.767 ms. Weaviate: Weaviate, 5.700 ms. Milvus: Milvus, 2.367 ms

  • Antfly
  • Other systems
Details

Compared with v0.2.0, Antfly’s 50K serial p95, peak throughput, and memory demand are less favorable. The tables show these results alongside recall.

Tested versions

Antfly
v0.2.1
Chroma
1.5.9
Elasticsearch
9.5.2
Postgres + pgvector
0.8.6 (PostgreSQL 16)
Weaviate
1.39.2
Milvus
3.0.0
Neo4j
5.26.29

50K quality and resources

Engine
Antfly
Workload
1536d-50K
Recall
0.9580
nDCG
0.9630
Peak QPS (queries/s)
984.7
p95 (ms)
3.533
Demand working set (MB)
6,992.3 MB
Cache-inclusive peak (MB)
7,746.0 MB
Final database volume (MB)
354.8 MB
Trials
3
Engine
Chroma
Workload
1536d-50K
Recall
0.9870
nDCG
0.9890
Peak QPS (queries/s)
180.9
p95 (ms)
8.067
Demand working set (MB)
463.3 MB
Cache-inclusive peak (MB)
812.9 MB
Final database volume (MB)
332.5 MB
Trials
3
Engine
Elasticsearch
Workload
1536d-50K
Recall
0.9560
nDCG
0.9640
Peak QPS (queries/s)
4,434.6
p95 (ms)
2.033
Demand working set (MB)
1,941.0 MB
Cache-inclusive peak (MB)
4,385.1 MB
Final database volume (MB)
312.1 MB
Trials
3
Engine
Postgres + pgvector
Workload
1536d-50K
Recall
0.9690
nDCG
0.9750
Peak QPS (queries/s)
2,721.0
p95 (ms)
4.167
Demand working set (MB)
2,303.7 MB
Cache-inclusive peak (MB)
3,533.5 MB
Final database volume (MB)
1,524.0 MB
Trials
3
Engine
Weaviate
Workload
1536d-50K
Recall
0.9550
nDCG
0.9630
Peak QPS (queries/s)
1,008.8
p95 (ms)
6.300
Demand working set (MB)
953.8 MB
Cache-inclusive peak (MB)
1,511.3 MB
Final database volume (MB)
391.5 MB
Trials
3
Engine
Milvus
Workload
1536d-50K
Recall
0.9630
nDCG
0.9700
Peak QPS (queries/s)
5,019.7
p95 (ms)
1.700
Demand working set (MB)
2,819.0 MB
Cache-inclusive peak (MB)
5,557.7 MB
Final database volume (MB)
2,489.2 MB
Trials
3
Engine
Neo4j
Workload
1536d-50K
Recall
0.99657 [0.99640–0.99670]
nDCG
0.99680 [0.99670–0.99690]
Peak QPS (queries/s)
343.26 [341.53–345.52]
p95 (ms)
28.533 [28.200–28.700]
Demand working set (MB)
2362.9 MB [2358.8–2366.8]
Cache-inclusive peak (MB)
Not reported
Final database volume (MB)
3139.1 MB [3139.1–3139.1]
Trials
3

1M quality and resources

Engine
Antfly
Workload
768d-1M
Recall
0.9630
nDCG
0.9670
Peak QPS (queries/s)
86.5
p95 (ms)
14.900
Demand working set (MB)
15,732.8 MB
Cache-inclusive peak (MB)
18,866.7 MB
Final database volume (MB)
3,831.6 MB (provisional)
Trials
3
Clean stops / total
0 / 3
Engine
Chroma
Workload
768d-1M
Recall
0.9560
nDCG
0.9620
Peak QPS (queries/s)
238.7
p95 (ms)
6.600
Demand working set (MB)
3,693.6 MB
Cache-inclusive peak (MB)
7,289.8 MB
Final database volume (MB)
3,459.7 MB
Trials
3
Clean stops / total
3 / 3
Engine
Elasticsearch
Workload
768d-1M
Recall
0.9590
nDCG
0.9610
Peak QPS (queries/s)
3,110.8
p95 (ms)
3.233
Demand working set (MB)
2,095.2 MB
Cache-inclusive peak (MB)
11,426.8 MB
Final database volume (MB)
3,148.5 MB
Trials
3
Clean stops / total
3 / 3
Engine
Postgres + pgvector
Workload
768d-1M
Recall
0.9660
nDCG
0.9700
Peak QPS (queries/s)
955.8
p95 (ms)
10.767
Demand working set (MB)
2,317.3 MB
Cache-inclusive peak (MB)
11,974.9 MB
Final database volume (MB)
9,329.3 MB
Trials
3
Clean stops / total
3 / 3
Engine
Weaviate
Workload
768d-1M
Recall
0.9560
nDCG
0.9620
Peak QPS (queries/s)
1,300.8
p95 (ms)
5.700
Demand working set (MB)
7,345.1 MB
Cache-inclusive peak (MB)
13,791.8 MB
Final database volume (MB)
3,947.5 MB
Trials
3
Clean stops / total
3 / 3
Engine
Milvus
Workload
768d-1M
Recall
0.9610
nDCG
0.9670
Peak QPS (queries/s)
3,898.5
p95 (ms)
2.367
Demand working set (MB)
13,262.6 MB
Cache-inclusive peak (MB)
21,631.0 MB
Final database volume (MB)
33,718.6 MB
Trials
3
Clean stops / total
3 / 3

Tenant-scoped filtered ANN

Engine
Antfly
Recall@100
1.0000
nDCG
1.0000
Peak QPS (queries/s)
2,879.8
Peak p99 (ms)
47.57
Serial p99 (ms)
2.57
Demand (MB)
7,645.4
Final volume (MB)
356.5
Trials
3
Engine
Weaviate
Recall@100
1.0000
nDCG
1.0000
Peak QPS (queries/s)
1,143.9
Peak p99 (ms)
11.71
Serial p99 (ms)
6.73
Demand (MB)
962.6
Final volume (MB)
392.7
Trials
3
Engine
Elasticsearch
Recall@100
1.0000
nDCG
1.0000
Peak QPS (queries/s)
5,810.7
Peak p99 (ms)
13.51
Serial p99 (ms)
1.80
Demand (MB)
1,912.8
Final volume (MB)
312.4
Trials
3
Engine
Postgres + pgvector
Recall@100
0.9573
nDCG
0.9661
Peak QPS (queries/s)
167.6
Peak p99 (ms)
224.49
Serial p99 (ms)
56.07
Demand (MB)
2,303.7
Final volume (MB)
1,524.0
Trials
3
Engine
Chroma
Recall@100
0.9999
nDCG
0.9999
Peak QPS (queries/s)
111.6
Peak p99 (ms)
573.43
Serial p99 (ms)
79.97
Demand (MB)
499.9
Final volume (MB)
336.1
Trials
3

Methodology

  • Antfly v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449. Recorded host: Apple M4 Max (linux-arm64).
  • Three independently built indexes per workload. Search parameters are selected on calibration queries and frozen for the separate 800-query evaluation set. Memory is peak cgroup anonymous plus shared memory; database volume is post-stop allocated storage.
  • Neo4j’s 50K results are means of three fresh runs that passed the benchmark checks on Apple M4 Max / Podman 6.1.1, Linux ARM64. Each uses one HNSW index build (M=16, efConstruction=256, quantization disabled), 200 calibration queries and 800 evaluation queries. Final storage measurements passed the post-shutdown checks.
04

Inference

Gemma 4 E4B QAT generation · end-to-end output tok/s

Antfly: Antfly, 50.00 tok/s. llama-server: llama-server, 57.91 tok/s. Ollama: Ollama, 57.61 tok/s

  • Antfly
  • Other systems

GLiNER2 extraction throughput · Q4 vs FP32

Antfly: Antfly, 25.50 requests/s. GLiNER2: GLiNER2, 34.89 requests/s

  • Antfly
  • GLiNER2

Additional inference workload throughput

Additional inference workload throughput
WorkloadAntfly throughputReference runtimeReference throughput
Dense embedding8.28 ops/ssentence-transformers141.59 ops/s
Sparse embedding96.50 ops/sONNX Runtime253.38 ops/s
Reranking9.42 ops/ssentence-transformers91.24 ops/s
Chunking189.46 ops/stokenizers28,908.90 ops/s
Transcription0.63 ops/swhisper.cpp8.61 ops/s
Details

Antfly runs GLiNER2 Q4 through a local HTTP endpoint with a Metal backend; standalone GLiNER2 runs FP32 in-process. Each runtime uses three fresh runs and five timed requests per run, in alternating order on Apple M4 Max / native macOS ARM64. Both outputs contained the expected values. Antfly returned the name separately from the organization/country record; standalone GLiNER2 returned one complete person record.

Gemma 4 E4B QAT uses the same Q4_0 GGUF weights across Antfly v0.2.1, llama-server, and Ollama on Apple M4 Max / native macOS ARM64. Each runtime uses three fresh interleaved runs, three disjoint warmups, and the same 18 held-out prompts at concurrency one.

End-to-end output throughput is pooled visible output tokens divided by full request time. The common GGUF tokenizer excludes EOS and control tokens. The rate includes every response; 8 of 18 outputs met the strict typed-JSON criteria in every run for all three runtimes, with zero request errors.

Persistent HTTP services use fresh TCP connections per request, greedy decoding, a 131,072-token context limit, and prompt caching disabled or reset. The comparison covers 162 measured requests.

Gemma model revision 4b4a2c1d584be7264f87aac328a1bc739ce81b6c; model SHA-256 676c35070db6dbe52f93e9c864ee0fba4eddea94b9c875d9cb10daff453fbaee.

Tested versions

Antfly
v0.2.1 (e5261e9e)
llama-server
0.4.0 / b10809 (5266f24da)
Ollama
0.33.3
onnxruntime
1.28.0
sentence-transformers
5.6.1
tokenizers
4.45.2
GLiNER2
1.3.2
transformers
4.45.2
faster-whisper
1.2.1
whisper.cpp
1.9.2

GLiNER2 warm latency and throughput

Engine
Antfly
Quantization
Q4
Warm p50 (ms)
39.14
Warm p95 (ms)
41.89
Mean requests/s
25.50
Per-run requests/s
25.97 / 26.03 / 24.50
Engine
GLiNER2
Quantization
FP32
Warm p50 (ms)
27.96
Warm p95 (ms)
30.95
Mean requests/s
34.89
Per-run requests/s
35.94 / 35.35 / 33.39

Gemma 4 E4B QAT serving results

Runtime
Antfly
Outputs meeting criteria per run
8/18 · 8/18 · 8/18
End-to-end output tok/s
50.00 [49.86–50.15]
Aggregate output tok/s
49.99 [49.85–50.14]
TTFT p50 (ms)
429.1 [428.8–429.5]
Post-first-token tok/s
72.33 [72.05–72.63]
Launch to ready (s)
0.732 [0.713–0.766]
Runtime
llama-server
Outputs meeting criteria per run
8/18 · 8/18 · 8/18
End-to-end output tok/s
57.91 [57.73–58.02]
Aggregate output tok/s
57.90 [57.72–58.01]
TTFT p50 (ms)
366.0 [365.8–366.1]
Post-first-token tok/s
83.03 [82.97–83.10]
Launch to ready (s)
5.354 [0.976–14.055]
Runtime
Ollama
Outputs meeting criteria per run
8/18 · 8/18 · 8/18
End-to-end output tok/s
57.61 [57.48–57.69]
Aggregate output tok/s
57.08 [56.95–57.16]
TTFT p50 (ms)
368.8 [368.4–369.1]
Post-first-token tok/s
82.53 [82.53–82.54]
Launch to ready (s)
1.269 [1.233–1.289]

Methodology

  • Each provider/model cell uses three fresh runs with one warmup per run.
  • Model artifacts use the format supported by each runtime; quantization and artifact identities are recorded with the results.
  • GLiNER2 reference model: fastino/gliner2-base-v1@8437ba583a733d87f56ae902f3b197934eedd58e. Antfly official v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449, binary SHA-256 e9b8bf12365508707568277601ba378ccbb52d760895d4f3ec06baf37be18fde.
05

PDF processing

Controlled OHR retrieval scores all 8,498 questions per input. Antfly v0.2.1 whole-PDF extraction scores 44.33% mean word-LCS; Elasticsearch’s bundled attachment processor with externally page-separated PDFs scores 47.49%. Matched ground-truth text scores 70.44%.

Controlled extraction · mean top-2 word-LCS

Input preparation differs: Antfly whole PDFs; Elasticsearch externally separated pages · common external retrieval evaluator

Input preparation differs: Antfly whole PDFs; Elasticsearch externally separated pages · common external retrieval evaluator. Antfly: Antfly, 44.33%. Matched ground truth: Matched ground truth, 70.44%. Elasticsearch: Elasticsearch, 47.49%

  • Antfly
  • Other systems

Controlled extraction · evidence scores (%)

Document source
Antfly v0.2.1 extraction
TXT
58.79
TAB
45.24
FOR
36.90
CHA
23.44
RO
6.89
MUL
40.88
ALL (%)
44.33
Extraction outcome
73 / 8,561 (0.85%) OCR failures
Document source
Matched OHR ground truth
TXT
81.75
TAB
70.10
FOR
75.91
CHA
68.24
RO
10.05
MUL
62.68
ALL (%)
70.44
Extraction outcome
Document source
Elasticsearch 9.5.2 Basic · externally page-prepared
TXT
62.70
TAB
51.84
FOR
33.79
CHA
25.70
RO
6.51
MUL
47.55
ALL (%)
47.49
Extraction outcome
1,200 empty pages; 0 processor request failures

Antfly whole-PDF extraction outcomes

Outcome
PDF loading
Valid / total
1,261 / 1,261
Observed result
0 load failures
Failure rate (%)
0.00%
Outcome
Page retention
Valid / total
8,561 / 8,561
Observed result
0 dropped pages
Failure rate (%)
0.00%
Outcome
Pages without OCR failure
Valid / total
8,488 / 8,561
Observed result
73 failures across 35 PDFs
Failure rate (%)
0.85%
Outcome
Retrievable chunks
Valid / total
1,259 / 1,261
Observed result
2 zero-chunk PDFs
Failure rate (%)
0.16%
Details

Elasticsearch 9.5.2 Basic uses pypdf for structural page separation and its bundled attachment processor for text extraction. Page identity comes from the prepared inputs. All 1,261 PDFs / 8,561 pages are retained, including 1,200 empty pages and zero processor request failures. A common external evaluator scores all 8,498 questions with both retrieval methods, including 42 strict filename-alias mismatches.

Tested versions

Antfly
v0.2.1
Matched ground truth
OHR-Bench 1f421eb428f9f5b8ac0bc8064d6ad1f13fab7af7 · matched text input
Elasticsearch
9.5.2

Published OHR reference scores · separate protocol (%)

Evidence group
Published upstream table
Document source
Ground Truth
TXT
81.6
TAB
69.8
FOR
75.2
CHA
70.3
RO
9.8
MUL
ALL (%)
70.4
OCR-failed pages
Evidence group
Published upstream table
Document source
MinerU 0.9.3
TXT
68.1
TAB
48.6
FOR
51.3
CHA
16.5
RO
5.9
MUL
ALL (%)
50.5
OCR-failed pages
Evidence group
Published upstream table
Document source
Marker 1.2.3
TXT
75.5
TAB
58.2
FOR
55.5
CHA
20.0
RO
5.9
MUL
ALL (%)
57.0
OCR-failed pages
Evidence group
Published upstream table
Document source
Azure Document Intelligence
TXT
78.0
TAB
59.4
FOR
55.2
CHA
45.2
RO
5.8
MUL
ALL (%)
60.6
OCR-failed pages
Evidence group
Published upstream table
Document source
GOT
TXT
62.5
TAB
41.1
FOR
49.0
CHA
17.4
RO
3.7
MUL
ALL (%)
45.8
OCR-failed pages
Evidence group
Published upstream table
Document source
Nougat
TXT
59.5
TAB
32.8
FOR
44.3
CHA
11.3
RO
4.4
MUL
ALL (%)
41.2
OCR-failed pages
Evidence group
Published upstream table
Document source
Qwen2.5-VL-72B
TXT
75.1
TAB
60.0
FOR
60.0
CHA
38.2
RO
5.3
MUL
ALL (%)
59.6
OCR-failed pages
Evidence group
Published upstream table
Document source
InternVL2.5-78B
TXT
68.6
TAB
57.9
FOR
55.6
CHA
45.1
RO
2.7
MUL
ALL (%)
56.2
OCR-failed pages
Evidence group
Published upstream table
Document source
olmOCR-7B-0225-preview
TXT
72.5
TAB
58.4
FOR
55.4
CHA
24.8
RO
5.0
MUL
ALL (%)
56.6
OCR-failed pages
Evidence group
Published upstream table
Document source
MonkeyOCR
TXT
74.6
TAB
56.5
FOR
55.5
CHA
16.5
RO
5.7
MUL
ALL (%)
55.9
OCR-failed pages
Evidence group
Published upstream table
Document source
Nanonets-OCR-s
TXT
71.8
TAB
59.8
FOR
57.4
CHA
43.7
RO
4.4
MUL
ALL (%)
58.3
OCR-failed pages

Elasticsearch page-prepared PDF input · matched retrieval methods

Text input
Elasticsearch attachment · page-prepared
BM25 word-LCS (%)
48.5495
BGE-M3 word-LCS (%)
46.4327
Mean word-LCS (%)
47.4911
Questions per method
8498
Text input
Matched ground-truth text
BM25 word-LCS (%)
71.3427
BGE-M3 word-LCS (%)
69.5429
Mean word-LCS (%)
70.4428
Questions per method
8498

Methodology

  • Antfly v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449. Recorded host: Apple M5 Pro (darwin-arm64).
  • Antfly's whole-PDF text and matched ground-truth text use the same 1,261-document candidate manifest, canonical logical paths, tokenizer/model sources, 1,024-token chunks with zero overlap, and deterministic UUID5 node IDs. The score averages word-level evidence LCS across all 8,498 questions and two independent top-2 retrieval methods: BM25 and exact full-precision BGE-M3. Unavailable evidence contributes zero.
  • Antfly retained all 8,561 pages. OCR failed on 73 pages across 35 PDFs (0.85%); these pages can still contain native text. Two PDFs produced zero retrievable chunks. These outcomes did not meet the benchmark’s structural checks; the controlled retrieval score includes every question.
  • The controlled corpus contains 1,261 PDFs. Canonical paths and UUID5 IDs fix document identity; 42 filename-mismatch questions receive zero under literal upstream scoring. The published reference table uses the upstream corpus and identity controls. Timing records include concurrent workloads.
  • The Elasticsearch extraction and matched ground-truth control use Apple M4 Max, a frozen candidate manifest, canonical logical paths, deterministic node IDs, 1,024-token chunks with zero overlap, top-2 BM25 and exact full-precision BGE-M3, MPS and batch size 8. ALL is the per-question mean across both independent methods. Matched ground truth scores 70.4428338881%.
06

Graph

How fast is a filtered relationship lookup? The LSQB-derived F2 workload follows Forum → Post → Tag across 432,235 nodes and 1,827,523 relationships, returning the matching Post IDs. Compared with v0.2.0, measured F2 p95 falls from 2.765 to 0.530 ms (80.8% lower).

Filtered two-hop lookup · serial p95

Serial p95 · in-VM and native macOS client runs; aggregation and transport in Details

Serial p95 · in-VM and native macOS client runs; aggregation and transport in Details. Antfly: Antfly, 0.530 ms. Neo4j: Neo4j, 0.868 ms. Elasticsearch: Elasticsearch, 1.310 ms. Postgres + pgvector: Postgres + pgvector, 0.294 ms. Weaviate: Weaviate, 1.340 ms

  • Antfly
  • Other systems

Filtered two-hop lookup · serving memory

Antfly: Antfly, 1,982.0 MB. Neo4j: Neo4j, 4,921.3 MB. Elasticsearch: Elasticsearch, 4990.4 MB. Postgres + pgvector: Postgres + pgvector, 150.2 MB. Weaviate: Weaviate, 1743.8 MB

  • Antfly
  • Other systems

Filtered two-hop lookup · database volume

Final database volumes that passed post-shutdown checks · Elasticsearch storage is provisional (see Details)

Final database volumes that passed post-shutdown checks · Elasticsearch storage is provisional (see Details). Antfly: Antfly, 200.7 MB. Neo4j: Neo4j, 635.5 MB. Postgres + pgvector: Postgres + pgvector, 362.4 MB. Weaviate: Weaviate, 738.9 MB

  • Antfly
  • Other systems
Details

Each PostgreSQL, Elasticsearch and Weaviate run retains 432,235 nodes / 1,827,523 relationships, passes 1,000 count-and-BLAKE3 correctness checks and serves 5,000 timed queries. Across three runs per engine: 9,000/9,000 checks, 45,000 timed queries and zero failed memory samples. PostgreSQL uses core node/edge tables and one SQL join. Elasticsearch uses a Forum→Post parent join plus Post→Tag relationship-ID membership. Weaviate uses native cross-reference filters and direct inverse edges.

Tested versions

Antfly
v0.2.1
Neo4j
5.26.29
Elasticsearch
9.5.2
Postgres + pgvector
16.15 (Debian 16.15-1.pgdg12+2)
Weaviate
1.39.2

Filtered two-hop results

Engine
Antfly
Version
Antfly 0.2.1 (Zig runtime)
Query
F2
p50 (ms)
0.360
p95 (ms)
0.530
p99 (ms)
0.742
Demand working set (MB)
1,982.0 MB
Final database volume (MB)
200.7 MB
Served samples
15,000
Engine
Neo4j
Version
5.26.29
Query
F2
p50 (ms)
0.467
p95 (ms)
0.868
p99 (ms)
1.297
Demand working set (MB)
4,921.3 MB
Final database volume (MB)
635.5 MB
Served samples
15,000
Engine
Elasticsearch
Version
9.5.2
Query
Native macOS client · Forum-to-Post native parent join; Post-to-Tag direct relationship-ID membership; all other typed adjacencies retained
p50 (ms)
0.9093 [0.8910–0.9338]
p95 (ms)
1.3100 [1.2890–1.3475]
p99 (ms)
1.5837 [1.5758–1.5958]
Demand working set (MB)
4990.41 MB [4985.25–4995.28]
Final database volume (MB)
40.30 MB [40.25–40.34] · provisional, 0/3 post-shutdown checks passed
Served samples
15000
Engine
Postgres + pgvector
Version
16.15 (Debian 16.15-1.pgdg12+2)
Query
Native macOS client · core PostgreSQL node/edge tables and one SQL join; no graph extension
p50 (ms)
0.1993 [0.1924–0.2045]
p95 (ms)
0.2942 [0.2905–0.2980]
p99 (ms)
0.3920 [0.3846–0.3997]
Demand working set (MB)
150.16 MB [150.15–150.16]
Final database volume (MB)
362.44 MB [362.36–362.49]
Served samples
15000
Engine
Weaviate
Version
1.39.2
Query
Native macOS client · native cross-reference filters from Post to Forum and Tag; explicit direct inverse CONTAINER_OF reference
p50 (ms)
0.9044 [0.8831–0.9285]
p95 (ms)
1.3400 [1.3092–1.3770]
p99 (ms)
2.1210 [1.9755–2.2876]
Demand working set (MB)
1743.81 MB [1724.76–1770.20]
Final database volume (MB)
738.89 MB [737.91–740.11]
Served samples
15000

Methodology

  • Antfly v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449. Recorded host: Apple M4 Max (linux-arm64).
  • Warm serial endpoint-filtered two-hop lookup on LSQB SF0.1. Three independent runs, 5,000 timed samples each, with count-and-BLAKE3 correctness checks. Antfly and Neo4j latency is the median run-level percentile.
  • PostgreSQL 16.15, Elasticsearch 9.5.2 Basic and Weaviate 1.39.2 use three fresh runs on Apple M4 Max: Linux ARM64 Podman services, 10 vCPU, 24 GiB, swap disabled. Persistent native macOS clients connect through loopback-published ports; Antfly and Neo4j use in-VM clients. PostgreSQL, Elasticsearch and Weaviate figures are means of run-level percentiles, with ranges in the table.
07

Running it

Antfly’s compressed Linux arm64 server archive is 47.4 MB and contains a static executable. Other figures are compressed container-image payloads.

Linux arm64 deployment payload

Engine
Antfly
Format
Static executable
Download source
Antfly release archive
Download (MB)
47.4
Executable (MB)
83.6
Status
Available
Engine
Weaviate
Format
Weaviate server container image
Download source
OCI registry manifest
Download (MB)
88.2
Executable (MB)
Status
Available
Engine
Postgres + pgvector
Format
PostgreSQL + pgvector container image
Download source
OCI registry manifest
Download (MB)
154.1
Executable (MB)
Status
Available
Engine
Chroma
Format
Chroma server container image
Download source
OCI registry manifest
Download (MB)
165.2
Executable (MB)
Status
Available
Engine
Elasticsearch
Format
Container image with JVM
Download source
OCI registry manifest
Download (MB)
758.8
Executable (MB)
Status
Available
Details

Tested versions

Antfly
v0.2.1
Weaviate
1.39.2
Postgres + pgvector
0.8.6
Chroma
1.5.9
Elasticsearch
9.5.2

Methodology

  • Antfly v0.2.1, commit e5261e9e2937d03943b518974bf8063351ad1449. Recorded host: Apple M4 Max (linux-arm64).
  • Antfly is the byte size of Antfly’s release archive. Other systems’ payloads sum compressed layers in their pinned Linux arm64 image manifests. Registry overhead and expanded filesystem sizes are excluded. MB denotes decimal megabytes.
Try it yourself

Search 10,000 Wikipedia articles

Follow the quickstart to load the dataset and start searching with Antfly.

Open the quickstart →