Public benchmark

Code-retrieval quality and cost

Measured with v1.2.16 (2026-09-15).

Profile public-core: 1000 regression queries across 4 public CoIR tasks, repeated 3 times. No private corpus or local path is included.

Corpus sampling: 1 of 4 corpora sampled to 5,000 documents (qrel documents retained, remainder dropped deterministically). Full CoIR corpora: codefeedback-st 156,526 documents (5,000 indexed). Indexed in full: codetrans-dl (816), codetrans-contest (1,008), cosqa (20,604). Scores on sampled corpora are not comparable to full-corpus CoIR leaderboard numbers.

1000regression queries
4public tasks
50languages
3repetitions
1/4corpora sampled

Aggregate results

ModenDCG@10MRR@10P@5R@20Warm p95Index size
lexical0.32450.27130.07380.5600345.08 ms36.89 MiB
hash0.32650.27380.07370.5620212.75 ms69.73 MiB
hybrid0.32690.27440.07370.5623204.81 ms69.72 MiB
blended0.32500.27240.07290.5600206.47 ms106.05 MiB
neural0.32660.27520.07350.5580212.41 ms106.06 MiB

Population variance, phase timings, peak RSS, binary identity, dataset revisions, and checksums are retained in the raw JSON.

Change from frozen baseline

hybrid improves nDCG@10 by +40.68% and MRR@10 by +42.60% over hash at commit 49b1571de77a. The raw JSON retains every task and run.

Per-task quality

TaskModenDCG@10MRR@10R@20
codetrans-dllexical0.31750.20520.7556
codetrans-dlhash0.32140.20780.7611
codetrans-dlhybrid0.31970.20670.7611
codetrans-dlblended0.31680.20230.7722
codetrans-dlneural0.31740.20310.7722
codetrans-contestlexical0.55780.50810.7240
codetrans-contesthash0.56680.52160.7330
codetrans-contesthybrid0.56860.52350.7330
codetrans-contestblended0.56770.52240.7330
codetrans-contestneural0.57030.52840.7285
cosqalexical0.14590.10740.3660
cosqahash0.14590.10740.3640
cosqahybrid0.14590.10740.3647
cosqablended0.14370.10580.3560
cosqaneural0.14160.10420.3520
codefeedback-stlexical0.71890.69030.8182
codefeedback-sthash0.71170.68060.8182
codefeedback-sthybrid0.71440.68430.8182
codefeedback-stblended0.71310.68260.8182
codefeedback-stneural0.73320.70510.8283

Every retained task remains visible so aggregate improvements cannot hide regressions.

Scope

Matrix covers natural-language and code-to-code retrieval. Exact-search tools require a separate exact-query workload.

Public-core is regression evidence, not an unseen-query generalization set: it overlaps the checkout-reference reranker's fit data. Runtime-reported model bytes match the checksum-bound fit ledger. Its query set and acceptance gates are unchanged.

Corpus sampling: 1 of 4 corpora sampled to 5,000 documents (qrel documents retained, remainder dropped deterministically). Full CoIR corpora: codefeedback-st 156,526 documents (5,000 indexed). Indexed in full: codetrans-dl (816), codetrans-contest (1,008), cosqa (20,604). Scores on sampled corpora are not comparable to full-corpus CoIR leaderboard numbers.

Artifacts use pinned public datasets and retain source revisions and checksums in raw JSON. Deterministic samples: codefeedback-st (99 queries, 5000 documents). Dataset cards do not declare licenses for: codetrans-dl, codetrans-contest, cosqa, codefeedback-st. Treat downloaded corpora as evaluation inputs; do not redistribute them without checking upstream terms.