Model Atlas · Decision theory

Markov Decision Process

A formal model of states, actions, transition probabilities, rewards and discounting.

core mathematical viewM=(S,A,P,R,γ)
Mental model

Understand it before memorizing it.

At each state choose an action then the environment moves and emits reward.

Best fit

Where this model earns its place

Sequential decision problems

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

RL modeling

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Policy analysis

Start with a simpler baseline then compare this model using the same evaluation split and operational constraints.

Strengths and limits

Trade-offs matter more than popularity.

Strengths

✓ Clear mathematical structure

✓ Supports planning

Limitations

△ State design can be difficult

△ Markov assumption may be imperfect

Evaluation

Metrics to watch

ReturnInterpret with the product objective and error cost.
Policy valueInterpret with the product objective and error cost.
State coverageInterpret with the product objective and error cost.
Production checklist

Before it reaches real users

  1. 01

    Define state boundaries

    Document the assumption and instrument the condition so regressions can be detected.

  2. 02

    Validate reward

    Document the assumption and instrument the condition so regressions can be detected.

  3. 03

    Test transition assumptions

    Document the assumption and instrument the condition so regressions can be detected.