module 4 of 7 · 50 min

A/B testing

Compare product variants with randomized experiments and explicit decision rules.

Explain A/B testing clearlyImplement a small A/B testing exampleEvaluate whether A/B testing improves a simpler baselineIdentify failure cases and operational constraints
learning statenot started
0% completesign in to track progress
mental model

start with the idea before the implementation.

Randomization aims to make groups comparable so outcome differences can be attributed to the intervention.
core concepts

the mechanisms you need to reason about.

01

Randomization

Randomization is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

02

Primary metric

Primary metric is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

03

Power

Power is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

04

Guardrails

Guardrails is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

05

Sequential monitoring

Sequential monitoring is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.

engineering lab

turn the lesson into evidence.

LAB 1

design an experiment

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 2

estimate sample size

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

LAB 3

write a decision memo

Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.

knowledge checks

prove you can explain and decide.

1

define unit of randomization

ask cortex to test me →
2

avoid peeking bias

ask cortex to test me →
3

check novelty effects

ask cortex to test me →
failure modes

what usually goes wrong.

risk

too many metrics

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

underpowered tests

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

risk

broken randomization

Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.

proof of learning

Create a short A/B testing engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.

Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.