Model Atlas · Information retrieval

BM25

A strong lexical ranking algorithm balancing term frequency, rarity and document length.

core mathematical viewΣ IDF(q) · f(k₁+1)/(f+k₁(1-b+b|D|/avgdl))
Mental model

Understand it before memorizing it.

Reward rare query terms that occur in a document while preventing repetition and long documents from dominating.

Best fit

Where this model earns its place

Search

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

RAG lexical retrieval

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Knowledge lookup

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Transparent

✓ Fast

✓ Excellent lexical baseline

Limitations

△ Misses semantic paraphrases

Evaluation

Metrics to watch

MRRInterpret with the product objective and error cost.
nDCG@kInterpret with the product objective and error cost.
Precision@kInterpret with the product objective and error cost.
Recall@kInterpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Tune k1/b

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Combine with semantic retrieval

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Keep evaluation judgments

    Document the assumption and instrument the condition so regressions can be detected.