Benchmarks

Measured, not claimed.

These are Hayfork's retrieval scores on public datasets, produced by the same code that serves every search, on the embedding profile the hosted service runs. The datasets are open, the method is described below, and the console shows the same numbers.

How to read the scores. Each query in a dataset comes with the documents a human judged relevant. nDCG@10 rewards putting relevant documents near the top of the first ten results and is the standard headline metric for retrieval. MRR@10 is how high the first relevant result sits. Recall@10 is the share of relevant documents found in the first ten. All three run from 0 to 1 and are shown as percentages.

NanoSciFact

Scientific claim retrieval. 2,919 documents, 50 queries, 56 relevance judgments. CC BY 4.0 dataset. Measured 3 October 2026 with the e5-base-v2 profile (768-dimensional vectors).

Hayfork hybridnDCG@10MRR@10Recall@10Hit rate
Result74.1%70.0%89.0%90.0%

System comparison

Same corpus, queries, judgments, chunking and top 10. Best ranking quality first. Hybrid search scores +10.4 nDCG points against keyword search on the same chunks.

SystemnDCG@10MRR@10Recall@10Hit rate
Hayfork hybrid
Hybrid BM25 + intfloat/e5-base-v2 vectors + RRF
74.1%70.0%89.0%90.0%
Keyword search (BM25)
Lexical ranking over the same chunks: what a keyword-only search engine returns out of the box
63.7%60.1%78.0%80.0%

Published reference points

Numbers reported by others on the same or the parent dataset. They are not measured by us and the setups differ, so they frame the result rather than rank it.

SystemnDCG@10Note
BM25 on NanoSciFact0.71First-stage BM25 baseline implied by the Sentence Transformers NanoBEIR reranker evaluation (reranked 0.7548, uplift +0.0449). Source.
e5-base-v2 alone, full SciFact0.72The embedding model Hayfork runs, as a pure vector retriever on the 5,183-document BEIR SciFact set (MTEB). Source.
BM25, full SciFact0.665Classic BM25 on the full BEIR SciFact set, from the E5 paper (Table 1). Source.

Method

What this does not show

Verifying the numbers

The datasets are public and the metrics are the standard ones used by the BEIR and MTEB retrieval benchmarks. If you want the full per-query results or help running the same measurement on your own content, write to us.