Benchmarks
Measured, not claimed.
These are Hayfork's retrieval scores on public datasets, produced by the same code that serves every search, on the embedding profile the hosted service runs. The datasets are open, the method is described below, and the console shows the same numbers.
NanoSciFact
Scientific claim retrieval. 2,919 documents, 50 queries, 56 relevance judgments.
CC BY 4.0 dataset. Measured 3 October 2026 with the e5-base-v2 profile (768-dimensional vectors).
| Hayfork hybrid | nDCG@10 | MRR@10 | Recall@10 | Hit rate |
|---|---|---|---|---|
| Result | 74.1% | 70.0% | 89.0% | 90.0% |
System comparison
Same corpus, queries, judgments, chunking and top 10. Best ranking quality first. Hybrid search scores +10.4 nDCG points against keyword search on the same chunks.
| System | nDCG@10 | MRR@10 | Recall@10 | Hit rate |
|---|---|---|---|---|
| Hayfork hybrid Hybrid BM25 + intfloat/e5-base-v2 vectors + RRF | 74.1% | 70.0% | 89.0% | 90.0% |
| Keyword search (BM25) Lexical ranking over the same chunks: what a keyword-only search engine returns out of the box | 63.7% | 60.1% | 78.0% | 80.0% |
Published reference points
Numbers reported by others on the same or the parent dataset. They are not measured by us and the setups differ, so they frame the result rather than rank it.
| System | nDCG@10 | Note |
|---|---|---|
| BM25 on NanoSciFact | 0.71 | First-stage BM25 baseline implied by the Sentence Transformers NanoBEIR reranker evaluation (reranked 0.7548, uplift +0.0449). Source. |
| e5-base-v2 alone, full SciFact | 0.72 | The embedding model Hayfork runs, as a pure vector retriever on the 5,183-document BEIR SciFact set (MTEB). Source. |
| BM25, full SciFact | 0.665 | Classic BM25 on the full BEIR SciFact set, from the E5 paper (Table 1). Source. |
Method
- Same pipeline as production. Documents are chunked with the production chunker, embedded with the passage prefix the model expects, and indexed into an isolated OpenSearch index. Queries are embedded with the query prefix and run through the same keyword plus vector fusion as
POST /search. - Document-level scoring. Hits are collapsed to documents before scoring, so a document that matches with three chunks counts once.
- Binary judgments. A document is relevant or it is not, as the dataset states. No graded relevance.
- No tuning on the test set. The retrieval settings are the hosted defaults. Nothing was adjusted to this dataset.
What this does not show
- One domain so far. Scientific claim retrieval is demanding, but your content is not scientific abstracts. The console's Benchmarks page lets you build the same measurement from your own queries and documents.
- Fifty queries is a small sample. Differences of a point or two between systems are within noise.
- Quality, not speed. This page reports ranking quality only. Response times depend on where the service runs and are not part of the published result.
Verifying the numbers
The datasets are public and the metrics are the standard ones used by the BEIR and MTEB retrieval benchmarks. If you want the full per-query results or help running the same measurement on your own content, write to us.