Benchmark

Cortrix retrieval benchmarks

Measured retrieval results for Cortrix semantic storage across four BEIR corpora, with source records and query scope.

August 2026 · Recall@10 and nDCG@10 · Public result bundle

Strongest measured result for each corpus

Four corpora. Four configurations. Each card shows its corpus’s highest measured Recall@10 and nDCG@10.

SciFactScientific claims
Recall@100.8752
nDCG@100.7784
Best measured armFull Stack

All 300 judged queries. This RC2 arm uses ingest-time LLM enrichment and query-time listwise reranking, with rag_fusion=false.

View SciFact source results
NFCorpusBiomedical retrieval
Recall@100.1768
nDCG@100.3772
Best measured armFull Stack

All 323 judged queries. This RC2 arm uses ingest-time LLM enrichment and query-time listwise reranking, with rag_fusion=false.

View NFCorpus source results
FiQAFinance QA
Recall@100.4554
nDCG@100.3156
Best measured armEmbedding + Reranking (full profile)

All 648 judged queries. No LLM arm was measured for FiQA in this RC2 bundle.

View FiQA source results
QuoraDuplicate questions
Recall@100.7142
nDCG@100.5003
Best measured armDense + BM25, no reranking

The first 2,000 of 10,000 judged queries against all 522,931 documents. This selected result has no LLM or reranking stage.

View Quora source results
Recall@10nDCG@10

Different configurations, not a cross-corpus ranking. Retrieval only. Quora covers the first 2,000 of 10,000 judged queries.

How to read these results and configurations

The four winners do not share one profile, so this is not a uniform Full Stack comparison or a cross-release comparison.

The SciFact and NFCorpus winners carry the bundle label Full Stack, but they are not like-for-like successors to the July Full Stack configuration. FiQA and Quora are won by non-LLM arms. All values are measured at top_k=10; timing remains available in the source bundle but is not presented here as server performance.

FiQA carries a known score-fusion loss in the measured build. Quora's absolute values are below published dense-retrieval baselines. Read the published limitations before quoting these values.

Both metrics use a 0–1 scale but measure different things. Their bar lengths are not expected to match. Read the metric definitions.

Methodology

From public corpus
to measured retrieval.

Inspect the full corpus, selected queries, configuration, and source records behind every score. These are maintainer-reported measurements, not an independent rerun.

Explore dataset scope

What was tested

SciFact, FiQA, NFCorpus, and Quora cover scientific claims, finance questions, biomedical retrieval, and duplicate-question retrieval. They are evidence for these recorded tasks, not every domain or a global ranking of difficulty.

Selection logic

Full corpora, declared queries, no hidden extrapolation.

Every published result uses the complete public corpus. SciFact, NFCorpus, and FiQA evaluate every judged test query. Quora uses the first 2,000 of 10,000 judged queries against all 522,931 documents, so it is explicitly not a full-query-set result.

All 16 source-bundle runs completed every selected query with zero request failures and zero missing query text. The scores remain maintainer-reported measurements; checksums and validators establish file identity and contract consistency, not an independent rerun.

CorpusDocumentsSelected queries
SciFact5,183300 of 300
NFCorpus3,633323 of 323
FiQA57,638648 of 648
Quora522,9312,000 of 10,000
BEIR in plain English

A retrieval benchmark, not an end-to-end agent demo

BEIR-style evaluation asks a narrow question: when a query is issued against a known corpus, does the system retrieve relevant documents near the top? Every published result uses the full corpus and evaluates its declared query set. The page keeps that query scope attached to Recall@10 and NDCG@10.

How to read the metrics

Recall finds relevant material. NDCG rewards ranking it higher.

Recall@10 checks whether relevant material appears in the top ten results. NDCG@10 also rewards placing stronger relevant results higher in that top-ten list. This bundle measured only top_k=10; deeper cutoffs are absent rather than estimated.

Scope boundary

Retrieval quality only

This benchmark does not claim server performance, third-party product superiority, business outcome gains, memory quality, audit coverage, security properties, or end-to-end RAG answer quality.

Inspect the evidence before rerunning

Start with the immutable bundle, its limitations, and its checksums. After obtaining upstream BEIR inputs, capped runs can check installation, pipeline behavior, source-ID mapping, and metric generation. Capped scores are not comparable to these measurements.

Method and bundle details

The bundle files and checksums are independently inspectable. These remain maintainer-reported measurements; a completed independent public rerun and independent per-query recomputation are not yet published.

Evidence fieldRecorded value or source
Run date2026-08-28 to 2026-08-31.
Execution boundaryCPU-only measurements. Hardware and timing for each result are recorded in the scorecards; this page does not present them as server-performance evidence.
Source rolesMeasured Core 79a4eb1; public runner 9490520; immutable result bundle 4b94390; RC2 Benchmark evidence carrier 5381a6a50795d546cf5b6dfc52fc892293033bf75381a6a . The bundle records the uncommitted measuring runners by content hash and does not mislabel the public rewrite as the measuring revision.
Dataset scopeFull SciFact, NFCorpus, FiQA, and Quora corpora. All judged queries are selected for the first three; Quora selects the first 2,000 of 10,000 judged queries.
MethodologyREADME, public runner, methodology, and reproduction guide.
Result bundleThe exact-commit bundle contains a manifest, summaries, 16 scorecards, reproduction instructions, and a complete SHA-256 inventory.
ScopeRetrieval accuracy at top_k=10 and inspectable artifact identity. No deeper cutoff, answer quality, server performance, security, competitive ranking, or business outcome is inferred.