Definition
BM25 is a lexical ranking function that scores a document against a query using term statistics. The score sums contributions from matching query terms.
Inverse document frequency gives more weight to terms that appear in fewer documents across the collection. Term-frequency saturation limits the benefit of repetition: each additional occurrence contributes less than the previous one. Document-length normalization adjusts the contribution using the document’s length relative to the collection’s average length.
Like TF-IDF, BM25 uses corpus statistics. Compared with a basic TF-IDF formula using raw term counts, BM25 builds saturation and document-length normalization into each term’s contribution. Settings control the strength of these effects.
Simple example
Suppose a query contains only “deadlock”, which appears in a small share of the incident reports. The collection has an average length of 200 tokens:
| Report | Length in tokens | Occurrences of “deadlock” |
|---|---|---|
| A | 200 | 3 |
| B | 1,000 | 3 |
| C | 200 | 6 |
The IDF factor is the same for all three reports because they match the same query term in the same collection.
With saturation and length normalization enabled, A gets a larger term contribution than B. Both mention “deadlock” three times. B is much longer than the collection average, so length normalization reduces its term contribution more. C gets a larger contribution than A, but doubling the occurrences does not double the contribution.
Why it matters
BM25 gives sparse retrieval a scoring rule that accounts for repetition and document length. It can rank passages containing technical terms or identifiers for a RAG system without requiring a neural model.
It still needs overlap in the analyzed terms. Different wording can leave a relevant passage without a matching term.
One important nuance
A BM25 score is a ranking value, not a calibrated probability of relevance. Its magnitude depends on the query, corpus statistics, and scoring settings. A score of 5 for one query may not mean the same thing as 5 for another.
Use BM25 to rank lexical matches for a given query. Do not treat a fixed score threshold as a universal test that a passage contains useful evidence.