Model Atlas · Reinforcement learning

Contextual Bandit

Chooses among actions using context while learning from immediate reward.

core mathematical viewa*=argmaxₐ E[r|x,a]
Mental model

Understand it before memorizing it.

Given the current context choose one option then learn from the immediate result.

Best fit

Where this model earns its place

Content selection

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Recommendations

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Experiment allocation

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Online adaptation

✓ Simpler than full RL

Limitations

△ No long-horizon credit assignment

Evaluation

Metrics to watch

RegretInterpret with the product objective and error cost.
RewardInterpret with the product objective and error cost.
Action coverageInterpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Log propensities

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Protect exploration

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Use guardrails

    Document the assumption and instrument the condition so regressions can be detected.