Assessment intelligence · Production foundation

Adaptive Interview Coach

A technical interview engine that samples from large domain banks, adapts difficulty, tracks weak concepts and schedules targeted review.

Weighted samplingBayesian-style mastery estimatesElo-inspired difficulty updatesSpaced repetitionContextual selection
7architecture stages
5models / algorithms
6data signals
4evaluation checks
4next milestones
01 · Problem definition

What this system is actually solving

Random question lists produce uneven practice and hide knowledge gaps. This project models mastery at topic level and deliberately revisits weak areas.

The project is designed around an observable outcome and explicit constraints. A production decision is only considered successful when the model or algorithm improves a baseline without creating unacceptable cost, latency, reliability or interpretability problems.

02 · Architecture

How the system is decomposed

  1. 01Question taxonomy
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  2. 02Difficulty model
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  3. 03Session sampler
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  4. 04Answer evaluator
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  5. 05Weak-topic state
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  6. 06Spaced-review queue
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  7. 07Readiness dashboard
    A separately testable stage with defined inputs, outputs, observability and failure handling.
03 · Models & algorithms

Why each algorithm exists

Weighted sampling

Role. Weighted sampling is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Bayesian-style mastery estimates

Role. Bayesian-style mastery estimates is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Elo-inspired difficulty updates

Role. Elo-inspired difficulty updates is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Spaced repetition

Role. Schedules review around memory strength so practice is concentrated where forgetting risk is highest.

Trade-off. A poor mastery estimate can schedule reviews too aggressively or too late.

Evaluate with. Recall after delay, review efficiency and mastery calibration.

Contextual selection

Role. Contextual selection is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

04 · Data design

Signals entering the system

Every useful model depends on the quality, timing and provenance of its inputs. CortexLab treats feature definitions and leakage checks as part of model engineering, not preprocessing trivia.

  • Correctness — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Response time — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Confidence — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Question difficulty — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Topic history — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Review intervals — captured with validation, lineage and monitoring so training and production meaning stay aligned.
05 · Evaluation

What has to be measured before calling it successful

  • Calibration of readiness
  • Weak-topic recovery
  • Retention after delay
  • Question coverage

Evaluation is split between offline quality, operational performance and failure analysis. A high headline metric does not override poor calibration, unstable segments, leakage or unusable latency.

06 · Failure analysis

Where this project can fail

Risk 1

Adaptive does not mean always harder

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 2

Coverage prevents tunnel vision

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 3

Readiness should be probabilistic

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

07 · Roadmap

How the project grows without becoming untestable

  1. 01Speech practice
  2. 02Company-style simulations
  3. 03Code execution
  4. 04Rubric-based explanations
08 · Interview lens

Questions an engineer should be able to answer

  • Why is this architecture preferable to a simpler baseline?
  • Which metric can look good while the product still fails?
  • Where can data leakage enter this pipeline?
  • What changes when traffic, data volume or latency requirements increase 10×?
  • Which part should be rolled back first if production quality drops?