module 2 of 7 · 40 min

MDPs

Model sequential decisions with states, actions, transitions, rewards and discounting.

Explain MDPs clearlyImplement a small MDPs exampleEvaluate whether MDPs improves a simpler baselineIdentify failure cases and operational constraints
learning statenot started
0% completesign in to track progress
mental model

start with the idea before the implementation.

The agent sees a state chooses an action then the environment transitions and returns reward.
core concepts

the mechanisms you need to reason about.

01

State

State is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

02

Action

Action is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

03

Transition

Transition is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

04

Reward and return

Reward and return is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

05

Policy

Policy is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

engineering lab

turn the lesson into evidence.

LAB 1

write a small MDP

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 2

draw transitions

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 3

calculate a return

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

knowledge checks

prove you can explain and decide.

1

test Markov assumption

ask cortex to test me →
2

define terminal states

ask cortex to test me →
3

explain discount factor

ask cortex to test me →
failure modes

what usually goes wrong.

risk

bad state representation

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

wrong reward

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

hidden history dependence

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

proof of learning

Create a short MDPs engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.

Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.