Innovation lab · Trainable NLP

Neural Intent Router

A planned small neural classifier that will replace parts of the rule router using locally trained examples while retaining deterministic fallbacks.

Bag-of-words baselineMLP classifierDistilled transformer laterTemperature calibration
6research stages
4algorithms under study
4evidence signals
4acceptance measures
4research milestones
Research question

The problem worth investigating

Regex intent routing is transparent and fast but brittle as language variety grows.

This innovation is treated as a falsifiable engineering hypothesis. The goal is not to prove that an advanced technique is impressive; it is to determine whether it produces a measurable improvement over a simpler control.

System proposal

Experimental architecture

  1. 01Training example store
    Instrumented independently so the experiment can reveal which stage contributes value.
  2. 02Tokenizer/features
    Instrumented independently so the experiment can reveal which stage contributes value.
  3. 03Small classifier
    Instrumented independently so the experiment can reveal which stage contributes value.
  4. 04Confidence calibration
    Instrumented independently so the experiment can reveal which stage contributes value.
  5. 05Rule fallback
    Instrumented independently so the experiment can reveal which stage contributes value.
  6. 06Evaluation dashboard
    Instrumented independently so the experiment can reveal which stage contributes value.
Methods

Algorithms under study

Bag-of-words baseline

Hypothesis. Bag-of-words baseline is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

MLP classifier

Hypothesis. MLP classifier is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Distilled transformer later

Hypothesis. Distilled transformer later is included because it addresses a specific measurable part of the system rather than being added as decoration.

Risk. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evidence. Task-specific quality metric, latency, reliability and failure-case analysis.

Temperature calibration

Hypothesis. Aligns predicted confidence with observed outcome frequency so thresholds mean something operationally.

Risk. Calibration can drift as the data distribution changes.

Evidence. Brier score, expected calibration error and reliability curves.

Evidence

What data the experiment needs

  • Labeled prompts — recorded with enough context to reproduce and audit the result.
  • Corrections — recorded with enough context to reproduce and audit the result.
  • Routing failures — recorded with enough context to reproduce and audit the result.
  • Synthetic paraphrases reviewed by humans — recorded with enough context to reproduce and audit the result.
Evaluation protocol

How CortexLab decides whether the idea survives

  • Macro F1
  • Calibration error
  • Latency
  • Fallback rate

Results should be compared against a control, segmented for failure cases and repeated across enough observations to avoid promoting noise into product behavior.

Threat model

What can go wrong

Wrong reward

Optimization can improve the metric while making the actual experience worse.

Overconfidence

Small or biased samples can make experimental gains look more certain than they are.

Distribution shift

A policy that works on past users may degrade as topics, traffic and behavior change.

Roadmap

Next research milestones

  1. 01Collect labels
  2. 02Train baseline
  3. 03A/B shadow routing
  4. 04Promote only after evaluation
Research notebook

Questions still open

  • Which simpler baseline must this beat before deployment?
  • How should uncertainty be calibrated and communicated?
  • What evidence would make us reject the idea?
  • How do we prevent reward hacking or accidental optimization of engagement alone?
  • Which decisions must remain human-reviewed?