Model Atlas · Information retrieval

TF-IDF

Represents text by upweighting terms frequent in a document but rare across the corpus.

core mathematical viewtfidf(t,d)=tf(t,d)·log(N/df(t))
Mental model

Understand it before memorizing it.

A term matters when it is important locally but not common everywhere.

Best fit

Where this model earns its place

Text similarity

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Search baseline

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Classical NLP

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Simple

✓ Interpretable

✓ Cheap

Limitations

△ Sparse

△ Weak semantics

Evaluation

Metrics to watch

Cosine similarityInterpret with the product objective and error cost.
Retrieval precisionInterpret with the product objective and error cost.
Classifier F1Interpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Normalize text consistently

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Watch vocabulary growth

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Use sparse matrices

    Document the assumption and instrument the condition so regressions can be detected.