AAI vs ZeroEntropy
eroEntropy's zembed-1 outperforms BAAI's bge-m3 on every aggregate metric we measure, including on long-tail domain queries (finance, legal, medical) where bge-m3's broader generalist training trails domain-tuned models. bge-m3 remains the strongest open-weight baseline in the public catalog and a sensible default for self-hosted retrieval starting points, particularly at its smaller parameter count (~568M vs zembed-1's 1.7B) when fitting onto a single commodity GPU is the binding constraint. The accuracy gap is the price you pay for that footprint advantage — and on the eval set, that price is consistent rather than situational.
- Smaller parameter count — bge-m3 is ~568M params vs zembed-1's 1.7B, so single-GPU self-hosting is lighter when accuracy headroom is not the priority.
- Aggregate NDCG@10 and Recall@100 across the eval set — zembed-1 leads on every vertical we measure.
- English instruction-following retrieval (Core17, News21, Robust04) where bge-m3's generalist training trails domain-tuned models.
- Domain-specific queries (finance, legal, medical) where specialized vocabulary penalizes bge-m3's broad training mix.
Three independent judges, one verdict.
Each judge scored every passage independently — different model, different temperature, different system prompt — and we report the NDCG@10 their relevance grades produce. BAAI's deltas don't depend on which judge you trust.
Headline deltas above are live from /evals/data/all-data.json;
narrative copy will be added when the head-to-head blog post is published.
Drilling into the per-vertical roll-up.
Per-vertical NDCG@10 deltas, averaged across the datasets in each vertical. Blue numbers = zembed-1 wins the vertical; warm-tan numbers = BAAI wins. Across 34 datasets head-to-head, zembed-1 wins on 33.
No cherry-picking. No hand-tuned splits.
Heterogeneous coverage — legal, finance, medical, multilingual, instruction-following, long-context. Every model evaluated on the same set.
gemini-3-flash, gpt-5-nano, grok-4-fast. Continuous 0–10 relevance scores; inter-judge agreement (κ) ≥ 0.7 across the suite. See eval-set-quality for the discipline.
Per-query deltas, not averaged independent samples. 95% CI on every reported number; statistical significance never asserted on n < 30.
All numbers on this page are sourced from /evals/. Latency figures use our open-source benchmark suite against public API endpoints.
