module 6 of 7 · 60 min

Q-learning

Learn action values off-policy from temporal-difference targets.

Explain Q-learning clearlyImplement a small Q-learning exampleEvaluate whether Q-learning improves a simpler baselineIdentify failure cases and operational constraints
learning statenot started
0% completesign in to track progress
mental model

start with the idea before the implementation.

Move the value of the chosen action toward immediate reward plus the best estimated next action value.
core concepts

the mechanisms you need to reason about.

01

Q values

Q values is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

02

TD target

TD target is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

03

Learning rate

Learning rate is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

04

Discount

Discount is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

05

Exploration

Exploration is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

engineering lab

turn the lesson into evidence.

LAB 1

implement tabular Q-learning

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 2

visualize Q-values

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 3

test epsilon schedules

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

knowledge checks

prove you can explain and decide.

1

derive the update

ask cortex to test me →
2

explain off-policy learning

ask cortex to test me →
3

measure policy stability

ask cortex to test me →
failure modes

what usually goes wrong.

risk

state explosion

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

reward sensitivity

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

unstable exploration

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

proof of learning

Create a short Q-learning engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.

Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.