The problem worth investigating
Automated SEO can over-optimize for clicks and damage trust or quality.
This innovation is treated as a falsifiable engineering hypothesis. The goal is not to prove that an advanced technique is impressive; it is to determine whether it produces a measurable improvement over a simpler control.
Experimental architecture
- 01Opportunity detector
Instrumented independently so the experiment can reveal which stage contributes value. - 02Candidate proposal
Instrumented independently so the experiment can reveal which stage contributes value. - 03Editorial approval
Instrumented independently so the experiment can reveal which stage contributes value. - 04Experiment assignment
Instrumented independently so the experiment can reveal which stage contributes value. - 05Outcome monitor
Instrumented independently so the experiment can reveal which stage contributes value. - 06Rollback
Instrumented independently so the experiment can reveal which stage contributes value.
Algorithms under study
Bandits
Hypothesis. Uses feedback to allocate more traffic to promising choices while reserving some exploration for alternatives.
Risk. Biased or sparse rewards can push the policy toward a locally attractive but globally poor choice.
Evidence. Reward lift, regret, exploration coverage and stability over time.
Decay detection
Hypothesis. Decay detection is included because it addresses a specific measurable part of the system rather than being added as decoration.
Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.
Semantic linking
Hypothesis. Semantic linking is included because it addresses a specific measurable part of the system rather than being added as decoration.
Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.
Query clustering
Hypothesis. Query clustering is included because it addresses a specific measurable part of the system rather than being added as decoration.
Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.
What data the experiment needs
- Search metrics — recorded with enough context to reproduce and audit the result.
- Article metadata — recorded with enough context to reproduce and audit the result.
- Engagement — recorded with enough context to reproduce and audit the result.
- Conversions — recorded with enough context to reproduce and audit the result.
How CortexLab decides whether the idea survives
- Incremental search value
- Quality guardrails
- Rollback rate
Results should be compared against a control, segmented for failure cases and repeated across enough observations to avoid promoting noise into product behavior.
What can go wrong
Wrong reward
Optimization can improve the metric while making the actual experience worse.
Overconfidence
Small or biased samples can make experimental gains look more certain than they are.
Distribution shift
A policy that works on past users may degrade as topics, traffic and behavior change.
Next research milestones
- 01Causal experiments
- 02Search-intent models
- 03Automated stale-code detection
Questions still open
- Which simpler baseline must this beat before deployment?
- How should uncertainty be calibrated and communicated?
- What evidence would make us reject the idea?
- How do we prevent reward hacking or accidental optimization of engagement alone?
- Which decisions must remain human-reviewed?