AAI vs ZeroEntropy

BAAI/ bge-m3
Datasets won0 / 34zembed-1 > bge-m3
NDCG@10 lead+0.0 ptsabsolute, averaged over 28 datasets
Recall@100 lead+0.0 ptsfirst-pass recall, top-100
per-dataset NDCG@10 delta · bar height = margin · blue = ZE · tan = BAAI

eroEntropy's zembed-1 outperforms BAAI's bge-m3 on every aggregate metric we measure, including on long-tail domain queries (finance, legal, medical) where bge-m3's broader generalist training trails domain-tuned models. bge-m3 remains the strongest open-weight baseline in the public catalog and a sensible default for self-hosted retrieval starting points, particularly at its smaller parameter count (~568M vs zembed-1's 1.7B) when fitting onto a single commodity GPU is the binding constraint. The accuracy gap is the price you pay for that footprint advantage — and on the eval set, that price is consistent rather than situational.

NDCG@10 lead +13.0 pts 0.721 vs 0.590
Recall@100 lead +12.9 pts 0.790 vs 0.661
Per-vertical Δ NDCG@10 (pts) sorted ZE-best → ZE-worst
Medical
+19.3
Instruction Following
+16.9
Specialized
+16.7
Finance
+15.9
Science
+14.7
Legal
+12.2
QA & Knowledge
+11.7
Manufacturing
+10.4
Multilingual
+8.5
BAAI wins ← → ZE wins
Where the gap closes
  • Smaller parameter count — bge-m3 is ~568M params vs zembed-1's 1.7B, so single-GPU self-hosting is lighter when accuracy headroom is not the priority.
Where ZE wins
  • Aggregate NDCG@10 and Recall@100 across the eval set — zembed-1 leads on every vertical we measure.
  • English instruction-following retrieval (Core17, News21, Robust04) where bge-m3's generalist training trails domain-tuned models.
  • Domain-specific queries (finance, legal, medical) where specialized vocabulary penalizes bge-m3's broad training mix.
Per-judge breakdown

Three independent judges, one verdict.

Gemini 3 Flashze leads
+0.0 pts NDCG@10
0 / 34 datasets won
zembed-1 0.728vsbge-m3 0.589
GPT-5 Nanoze leads
+0.0 pts NDCG@10
0 / 34 datasets won
zembed-1 0.734vsbge-m3 0.618
Grok 4 Fastze leads
+0.0 pts NDCG@10
0 / 34 datasets won
zembed-1 0.700vsbge-m3 0.563
Unanimous · all 3 judges picked zembed-1

Each judge scored every passage independently — different model, different temperature, different system prompt — and we report the NDCG@10 their relevance grades produce. BAAI's deltas don't depend on which judge you trust.

TL;DR · not yet authored

Headline deltas above are live from /evals/data/all-data.json; narrative copy will be added when the head-to-head blog post is published.

By vertical, by dataset

Drilling into the per-vertical roll-up.

Per-vertical NDCG@10 deltas, averaged across the datasets in each vertical. Blue numbers = zembed-1 wins the vertical; warm-tan numbers = BAAI wins. Across 34 datasets head-to-head, zembed-1 wins on 33.

How we measure

No cherry-picking. No hand-tuned splits.

28 datasets

Heterogeneous coverage — legal, finance, medical, multilingual, instruction-following, long-context. Every model evaluated on the same set.

3 LLM judges

gemini-3-flash, gpt-5-nano, grok-4-fast. Continuous 0–10 relevance scores; inter-judge agreement (κ) ≥ 0.7 across the suite. See eval-set-quality for the discipline.

Paired bootstrap

Per-query deltas, not averaged independent samples. 95% CI on every reported number; statistical significance never asserted on n < 30.

All numbers on this page are sourced from /evals/. Latency figures use our open-source benchmark suite against public API endpoints.

ZeroEntropy
The best AI teams build with ZeroEntropy models
Follow us on
GitHubTwitterSlackLinkedInDiscord