module 7 of 7 · 65 min

Policy gradients

Optimize a parameterized policy directly using sampled returns or advantages.

Explain Policy gradients clearlyImplement a small Policy gradients exampleEvaluate whether Policy gradients improves a simpler baselineIdentify failure cases and operational constraints
learning statenot started
0% completesign in to track progress
mental model

start with the idea before the implementation.

Increase the probability of actions that performed better than expected and reduce worse ones.
core concepts

the mechanisms you need to reason about.

01

Policy distribution

Policy distribution is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

02

Log-probability gradient

Log-probability gradient is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

03

Return

Return is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

04

Advantage

Advantage is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

05

Variance reduction

Variance reduction is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

engineering lab

turn the lesson into evidence.

LAB 1

derive REINFORCE intuition

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 2

simulate a policy update

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 3

compare with value methods

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

knowledge checks

prove you can explain and decide.

1

explain high variance

ask cortex to test me →
3

separate on-policy and off-policy

ask cortex to test me →
failure modes

what usually goes wrong.

risk

unstable rewards

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

poor exploration

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

high variance

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

proof of learning

Create a short Policy gradients engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.

Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.