start with the idea before the implementation.
the mechanisms you need to reason about.
Randomization
Randomization is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.
Primary metric
Primary metric is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.
Power
Power is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.
Guardrails
Guardrails is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.
Sequential monitoring
Sequential monitoring is studied through intuition implementation evidence and trade-offs. The goal is to be able to explain the mechanism and verify it with a concrete test rather than only repeat a definition.
turn the lesson into evidence.
design an experiment
Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.
estimate sample size
Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.
write a decision memo
Build the smallest version first. Record the input, expected output, measured result and one failure you discovered.
prove you can explain and decide.
define unit of randomization
ask cortex to test me →avoid peeking bias
ask cortex to test me →check novelty effects
ask cortex to test me →what usually goes wrong.
too many metrics
Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.
underpowered tests
Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.
broken randomization
Detect this early by defining a baseline, a measurable signal and a condition that would cause you to stop or redesign the approach.
Create a short A/B testing engineering note with one working artifact one metric one failure case and one decision about when you would or would not use it.
Save the result in your portfolio or project repository. A strong learning artifact should make your assumptions, metrics and failure analysis visible.