Model Atlas · Sequence models

Transformer

Uses attention to model relationships between tokens without recurrent state.

core mathematical viewAttention(Q,K,V)=softmax(QKᵀ/√d)V
Mental model

Understand it before memorizing it.

Each token dynamically decides which other tokens matter for its representation.

Best fit

Where this model earns its place

Language

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Multimodal sequences

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Long-range dependencies

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Parallel training

✓ Flexible context modeling

Limitations

△ Memory grows with context

△ Large data/compute needs

Evaluation

Metrics to watch

PerplexityInterpret with the product objective and error cost.
Task accuracyInterpret with the product objective and error cost.
LatencyInterpret with the product objective and error cost.
MemoryInterpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Control context size

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Cache wisely

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Evaluate hallucination/grounding

    Document the assumption and instrument the condition so regressions can be detected.