Understand it before memorizing it.
Update the value of the action you took toward the reward plus the best value available next.
Where this model earns its place
Discrete decision policies
Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.
Control problems
Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.
Policy experimentation
Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.
Trade-offs matter more than popularity.
Strengths
✓ Model-free
✓ Off-policy
Limitations
△ State/action explosion
△ Reward sensitivity
Metrics to watch
Before it reaches real users
- 01
Design reward carefully
Document the assumption and instrument the condition so regressions can be detected.
- 02
Keep safe exploration bounds
Document the assumption and instrument the condition so regressions can be detected.
- 03
Replay policies offline
Document the assumption and instrument the condition so regressions can be detected.