Adaptive education · Prototype family

Game-Based Learning Engine

A technical game platform where challenge selection responds to mastery, error patterns and pace rather than using empty engagement mechanics.

Dynamic difficulty adjustmentMastery scoringSpaced challenge schedulingBandit content selection
7architecture stages
4models / algorithms
6data signals
4evaluation checks
4next milestones
01 · Problem definition

What this system is actually solving

Educational games often optimize time-on-site instead of learning. This project treats mastery and transfer as the primary objectives.

The project is designed around an observable outcome and explicit constraints. A production decision is only considered successful when the model or algorithm improves a baseline without creating unacceptable cost, latency, reliability or interpretability problems.

02 · Architecture

How the system is decomposed

  1. 01Skill graph
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  2. 02Challenge library
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  3. 03Difficulty controller
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  4. 04Scoring service
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  5. 05Mastery estimator
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  6. 06Achievement system
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  7. 07Replay analytics
    A separately testable stage with defined inputs, outputs, observability and failure handling.
03 · Models & algorithms

Why each algorithm exists

Dynamic difficulty adjustment

Role. Dynamic difficulty adjustment is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Mastery scoring

Role. Mastery scoring is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Spaced challenge scheduling

Role. Spaced challenge scheduling is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Bandit content selection

Role. Uses feedback to allocate more traffic to promising choices while reserving some exploration for alternatives.

Trade-off. Biased or sparse rewards can push the policy toward a locally attractive but globally poor choice.

Evaluate with. Reward lift, regret, exploration coverage and stability over time.

04 · Data design

Signals entering the system

Every useful model depends on the quality, timing and provenance of its inputs. CortexLab treats feature definitions and leakage checks as part of model engineering, not preprocessing trivia.

  • Accuracy — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Latency — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Hints — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Retries — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Abandonment — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Topic mastery — captured with validation, lineage and monitoring so training and production meaning stay aligned.
05 · Evaluation

What has to be measured before calling it successful

  • Learning gain
  • Retention
  • Transfer to interview questions
  • Challenge completion

Evaluation is split between offline quality, operational performance and failure analysis. A high headline metric does not override poor calibration, unstable segments, leakage or unusable latency.

06 · Failure analysis

Where this project can fail

Risk 1

Addiction is not a learning metric

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 2

Failure needs useful feedback

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 3

Difficulty should stay in the productive struggle zone

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

07 · Roadmap

How the project grows without becoming untestable

  1. 01Multiplayer technical challenges
  2. 02Code sandbox
  3. 03Teacher dashboards
  4. 04Generative level builder from approved templates
08 · Interview lens

Questions an engineer should be able to answer

  • Why is this architecture preferable to a simpler baseline?
  • Which metric can look good while the product still fails?
  • Where can data leakage enter this pipeline?
  • What changes when traffic, data volume or latency requirements increase 10×?
  • Which part should be rolled back first if production quality drops?