All 300 judged queries. This RC2 arm uses ingest-time LLM enrichment and query-time listwise reranking, with rag_fusion=false.
Strongest measured result for each corpus
Four corpora. Four configurations. Each card shows its corpus’s highest measured Recall@10 and nDCG@10.
All 323 judged queries. This RC2 arm uses ingest-time LLM enrichment and query-time listwise reranking, with rag_fusion=false.
All 648 judged queries. No LLM arm was measured for FiQA in this RC2 bundle.
View FiQA source resultsThe first 2,000 of 10,000 judged queries against all 522,931 documents. This selected result has no LLM or reranking stage.
View Quora source resultsDifferent configurations, not a cross-corpus ranking. Retrieval only. Quora covers the first 2,000 of 10,000 judged queries.
How to read these results and configurations
The four winners do not share one profile, so this is not a uniform Full Stack comparison or a cross-release comparison.
The SciFact and NFCorpus winners carry the bundle label Full Stack, but they are not like-for-like successors to the July Full Stack configuration. FiQA and Quora are won by non-LLM arms. All values are measured at top_k=10; timing remains available in the source bundle but is not presented here as server performance.
FiQA carries a known score-fusion loss in the measured build. Quora's absolute values are below published dense-retrieval baselines. Read the published limitations before quoting these values.
Both metrics use a 0–1 scale but measure different things. Their bar lengths are not expected to match. Read the metric definitions.
From public corpus
to measured retrieval.
Inspect the full corpus, selected queries, configuration, and source records behind every score. These are maintainer-reported measurements, not an independent rerun.
Explore dataset scopeWhat was tested
SciFact, FiQA, NFCorpus, and Quora cover scientific claims, finance questions, biomedical retrieval, and duplicate-question retrieval. They are evidence for these recorded tasks, not every domain or a global ranking of difficulty.
Full corpora, declared queries, no hidden extrapolation.
Every published result uses the complete public corpus. SciFact, NFCorpus, and FiQA evaluate every judged test query. Quora uses the first 2,000 of 10,000 judged queries against all 522,931 documents, so it is explicitly not a full-query-set result.
All 16 source-bundle runs completed every selected query with zero request failures and zero missing query text. The scores remain maintainer-reported measurements; checksums and validators establish file identity and contract consistency, not an independent rerun.
| Corpus | Documents | Selected queries |
|---|---|---|
| SciFact | 5,183 | 300 of 300 |
| NFCorpus | 3,633 | 323 of 323 |
| FiQA | 57,638 | 648 of 648 |
| Quora | 522,931 | 2,000 of 10,000 |
A retrieval benchmark, not an end-to-end agent demo
BEIR-style evaluation asks a narrow question: when a query is issued against a known corpus, does the system retrieve relevant documents near the top? Every published result uses the full corpus and evaluates its declared query set. The page keeps that query scope attached to Recall@10 and NDCG@10.
Recall finds relevant material. NDCG rewards ranking it higher.
Recall@10 checks whether relevant material appears in the top ten results. NDCG@10 also rewards placing stronger relevant results higher in that top-ten list. This bundle measured only top_k=10; deeper cutoffs are absent rather than estimated.
Retrieval quality only
This benchmark does not claim server performance, third-party product superiority, business outcome gains, memory quality, audit coverage, security properties, or end-to-end RAG answer quality.
Inspect the evidence before rerunning
Start with the immutable bundle, its limitations, and its checksums. After obtaining upstream BEIR inputs, capped runs can check installation, pipeline behavior, source-ID mapping, and metric generation. Capped scores are not comparable to these measurements.
Method and bundle details
The bundle files and checksums are independently inspectable. These remain maintainer-reported measurements; a completed independent public rerun and independent per-query recomputation are not yet published.
| Evidence field | Recorded value or source |
|---|---|
| Run date | 2026-08-28 to 2026-08-31. |
| Execution boundary | CPU-only measurements. Hardware and timing for each result are recorded in the scorecards; this page does not present them as server-performance evidence. |
| Source roles | Measured Core 79a4eb1; public runner 9490520; immutable result bundle 4b94390; RC2 Benchmark evidence carrier 5381a6a50795d546cf5b6dfc52fc892293033bf75381a6a . The bundle records the uncommitted measuring runners by content hash and does not mislabel the public rewrite as the measuring revision. |
| Dataset scope | Full SciFact, NFCorpus, FiQA, and Quora corpora. All judged queries are selected for the first three; Quora selects the first 2,000 of 10,000 judged queries. |
| Methodology | README, public runner, methodology, and reproduction guide. |
| Result bundle | The exact-commit bundle contains a manifest, summaries, 16 scorecards, reproduction instructions, and a complete SHA-256 inventory. |
| Scope | Retrieval accuracy at top_k=10 and inspectable artifact identity. No deeper cutoff, answer quality, server performance, security, competitive ranking, or business outcome is inferred. |