Model Atlas · Reinforcement learning

Q-Learning

An off-policy temporal-difference method that learns action values from reward transitions.

core mathematical viewQ←Q+α[r+γ maxₐ Q(s′,a)-Q]
Mental model

Understand it before memorizing it.

Update the value of the action you took toward the reward plus the best value available next.

Best fit

Where this model earns its place

Discrete decision policies

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Control problems

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Policy experimentation

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Model-free

✓ Off-policy

Limitations

△ State/action explosion

△ Reward sensitivity

Evaluation

Metrics to watch

Average rewardInterpret with the product objective and error cost.
RegretInterpret with the product objective and error cost.
Policy stabilityInterpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Design reward carefully

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Keep safe exploration bounds

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Replay policies offline

    Document the assumption and instrument the condition so regressions can be detected.