Innovation lab · Interactive model improvement

Human Teaching Loop

When Cortex is uncertain, it asks the user to verify or teach the correct explanation. Approved corrections become reusable knowledge with provenance.

Confidence thresholdsSimilarity deduplicationCorrection scoringTrust-weighted retrieval
7research stages
4algorithms under study
5evidence signals
4acceptance measures
4research milestones
Research question

The problem worth investigating

Low-confidence systems either hallucinate or stop. This innovation makes uncertainty productive by converting it into a structured teaching event.

This innovation is treated as a falsifiable engineering hypothesis. The goal is not to prove that an advanced technique is impressive; it is to determine whether it produces a measurable improvement over a simpler control.

System proposal

Experimental architecture

  1. 01Confidence gate
    Instrumented independently so the experiment can reveal which stage contributes value.
  2. 02Clarification prompt
    Instrumented independently so the experiment can reveal which stage contributes value.
  3. 03Correction capture
    Instrumented independently so the experiment can reveal which stage contributes value.
  4. 04Normalization and deduplication
    Instrumented independently so the experiment can reveal which stage contributes value.
  5. 05Trust/provenance metadata
    Instrumented independently so the experiment can reveal which stage contributes value.
  6. 06Human review
    Instrumented independently so the experiment can reveal which stage contributes value.
  7. 07Knowledge promotion
    Instrumented independently so the experiment can reveal which stage contributes value.
Methods

Algorithms under study

Confidence thresholds

Hypothesis. Confidence thresholds is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Similarity deduplication

Hypothesis. Similarity deduplication is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Correction scoring

Hypothesis. Correction scoring is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Trust-weighted retrieval

Hypothesis. Trust-weighted retrieval is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Evidence

What data the experiment needs

  • User correction — recorded with enough context to reproduce and audit the result.
  • Original question — recorded with enough context to reproduce and audit the result.
  • Prior answer — recorded with enough context to reproduce and audit the result.
  • Feedback — recorded with enough context to reproduce and audit the result.
  • Reviewer approval — recorded with enough context to reproduce and audit the result.
Evaluation protocol

How CortexLab decides whether the idea survives

  • Correction acceptance rate
  • Repeat-question improvement
  • Bad-correction rejection
  • Knowledge coverage growth

Results should be compared against a control, segmented for failure cases and repeated across enough observations to avoid promoting noise into product behavior.

Threat model

What can go wrong

Wrong reward

Optimization can improve the metric while making the actual experience worse.

Overconfidence

Small or biased samples can make experimental gains look more certain than they are.

Distribution shift

A policy that works on past users may degrade as topics, traffic and behavior change.

Roadmap

Next research milestones

  1. 01Expert reputation
  2. 02Consensus teaching
  3. 03Contradiction detector
  4. 04Knowledge versioning
Research notebook

Questions still open

  • Which simpler baseline must this beat before deployment?
  • How should uncertainty be calibrated and communicated?
  • What evidence would make us reject the idea?
  • How do we prevent reward hacking or accidental optimization of engagement alone?
  • Which decisions must remain human-reviewed?