Innovation lab · Learning optimization

Adaptive Difficulty Controller

Chooses challenge difficulty from recent correctness, response time, hint usage and mastery estimates to maintain productive difficulty.

Elo-inspired updatesBandit selectionSpaced repetitionMastery thresholds
5research stages
4algorithms under study
5evidence signals
3acceptance measures
3research milestones
Research question

The problem worth investigating

Static difficulty bores advanced learners and overwhelms beginners.

This innovation is treated as a falsifiable engineering hypothesis. The goal is not to prove that an advanced technique is impressive; it is to determine whether it produces a measurable improvement over a simpler control.

System proposal

Experimental architecture

  1. 01Mastery state
    Instrumented independently so the experiment can reveal which stage contributes value.
  2. 02Challenge metadata
    Instrumented independently so the experiment can reveal which stage contributes value.
  3. 03Difficulty policy
    Instrumented independently so the experiment can reveal which stage contributes value.
  4. 04Reward model
    Instrumented independently so the experiment can reveal which stage contributes value.
  5. 05Review scheduler
    Instrumented independently so the experiment can reveal which stage contributes value.
Methods

Algorithms under study

Elo-inspired updates

Hypothesis. Maintains a compact estimate of learner skill and challenge difficulty that updates after each attempt.

Risk. Assumes a simplified relationship between ability and item difficulty.

Evidence. Prediction calibration, ranking accuracy and learning progression.

Bandit selection

Hypothesis. Uses feedback to allocate more traffic to promising choices while reserving some exploration for alternatives.

Risk. Biased or sparse rewards can push the policy toward a locally attractive but globally poor choice.

Evidence. Reward lift, regret, exploration coverage and stability over time.

Spaced repetition

Hypothesis. Schedules review around memory strength so practice is concentrated where forgetting risk is highest.

Risk. A poor mastery estimate can schedule reviews too aggressively or too late.

Evidence. Recall after delay, review efficiency and mastery calibration.

Mastery thresholds

Hypothesis. Mastery thresholds is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Evidence

What data the experiment needs

  • Answers — recorded with enough context to reproduce and audit the result.
  • Time — recorded with enough context to reproduce and audit the result.
  • Hints — recorded with enough context to reproduce and audit the result.
  • Retries — recorded with enough context to reproduce and audit the result.
  • Topic — recorded with enough context to reproduce and audit the result.
Evaluation protocol

How CortexLab decides whether the idea survives

  • Learning gain
  • Drop-off
  • Mastery calibration

Results should be compared against a control, segmented for failure cases and repeated across enough observations to avoid promoting noise into product behavior.

Threat model

What can go wrong

Wrong reward

Optimization can improve the metric while making the actual experience worse.

Overconfidence

Small or biased samples can make experimental gains look more certain than they are.

Distribution shift

A policy that works on past users may degrade as topics, traffic and behavior change.

Roadmap

Next research milestones

  1. 01Personalized pacing
  2. 02Cross-domain transfer
  3. 03Longitudinal learning models
Research notebook

Questions still open

  • Which simpler baseline must this beat before deployment?
  • How should uncertainty be calibrated and communicated?
  • What evidence would make us reject the idea?
  • How do we prevent reward hacking or accidental optimization of engagement alone?
  • Which decisions must remain human-reviewed?