module 1 of 7 · 35 min

Bandits

Learn which action performs best while continuing controlled exploration.

Explain Bandits clearlyImplement a small Bandits exampleEvaluate whether Bandits improves a simpler baselineIdentify failure cases and operational constraints
learning statenot started
0% completesign in to track progress
mental model

start with the idea before the implementation.

A bandit repeatedly chooses an action and immediately observes reward without modeling long-term state transitions.
core concepts

the mechanisms you need to reason about.

01

Exploration vs exploitation

Exploration vs exploitation is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

02

Epsilon strategies

Epsilon strategies is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

03

UCB

UCB is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

04

Thompson sampling

Thompson sampling is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

engineering lab

turn the lesson into evidence.

LAB 1

simulate two arms

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 2

plot cumulative regret

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 3

add guardrails

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

knowledge checks

prove you can explain and decide.

3

distinguish contextual bandits from MDPs

ask cortex to test me →
failure modes

what usually goes wrong.

risk

unsafe exploration

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

biased logging

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

reward hacking

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

proof of learning

Create a short Bandits engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.

Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.