Cortex Academy · Advanced

Reinforcement Learning & Decision Systems

Learn MDPs, value methods, policy learning and safe decision optimization.

BanditsMDPsDynamic programmingMonte CarloTD learningQ-learningPolicy gradients
ask cortex to assess me
modules7
projects3
assessment lenses4
your progress0%
Learning outcomes

What you should be able to do

1

Model states/actions/rewards

2

Implement value updates

3

Understand exploration

4

Evaluate policies offline

Curriculum

Every module is now a full lesson.

Applied work

Projects that prove the skill

Response policy lab

Use a baseline first then instrument the result and document failure cases.

design with cortex →

Adaptive difficulty controller

Use a baseline first then instrument the result and document failure cases.

design with cortex →

SEO bandit

Use a baseline first then instrument the result and document failure cases.

design with cortex →
Assessment lens

How progress is judged

Reward design

Evidence should come from code, results, explanation quality and the ability to identify when an approach should not be used.

Policy stability

Evidence should come from code, results, explanation quality and the ability to identify when an approach should not be used.

Regret

Evidence should come from code, results, explanation quality and the ability to identify when an approach should not be used.

Off-policy reasoning

Evidence should come from code, results, explanation quality and the ability to identify when an approach should not be used.

Turn the track into a personal plan.

Cortex can break this path into daily sessions and adjust based on completed lessons, interview results and your saved learning goals.

Ask Cortex to plan it →