First-party AI architecture · Active R&D

Cortex Local Intelligence Engine

The local NLP, retrieval, confidence, MDP-state and reinforcement-feedback engine powering Ask Cortex without an external generative-model API.

TF-IDFCosine similarityBM25Rule/feature intent routingContextual policy selectionTemporal-difference Q-learningConfidence calibration
10architecture stages
7models / algorithms
5data signals
6evaluation checks
6next milestones
01 · Problem definition

What this system is actually solving

Most AI products are wrappers around remote foundation-model APIs. CortexLab explores how much useful intelligence can be built from owned knowledge, retrieval, policy selection, feedback and progressively trained local models.

The project is designed around an observable outcome and explicit constraints. A production decision is only considered successful when the model or algorithm improves a baseline without creating unacceptable cost, latency, reliability or interpretability problems.

02 · Architecture

How the system is decomposed

  1. 01Unicode-safe text normalization
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  2. 02Intent and dialogue-act classifier
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  3. 03TF-IDF vector space
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  4. 04BM25 lexical retrieval
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  5. 05Knowledge graph/context scopes
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  6. 06Confidence calibration
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  7. 07MDP response state
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  8. 08Q-value response policy
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  9. 09Human correction memory
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  10. 10Conversation persistence
    A separately testable stage with defined inputs, outputs, observability and failure handling.
03 · Models & algorithms

Why each algorithm exists

TF-IDF

Role. Creates a transparent lexical representation that works well for technical terms and can be inspected directly.

Trade-off. It does not understand deep semantic equivalence when two passages use very different vocabulary.

Evaluate with. Retrieval precision@k, recall@k and query-level error analysis.

Cosine similarity

Role. Compares vector direction rather than raw magnitude, making it useful for normalized text representations.

Trade-off. Its quality is limited by the representation placed into the vector space.

Evaluate with. Pairwise relevance accuracy and retrieval ranking quality.

BM25

Role. Ranks documents by term relevance while controlling for document length and repeated terms.

Trade-off. Still lexical; it misses semantic matches that do not share useful vocabulary.

Evaluate with. MRR, nDCG@k, precision@k and judged relevance.

Rule/feature intent routing

Role. Rule/feature intent routing is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Contextual policy selection

Role. Contextual policy selection is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Temporal-difference Q-learning

Role. Temporal-difference Q-learning is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Confidence calibration

Role. Aligns predicted confidence with observed outcome frequency so thresholds mean something operationally.

Trade-off. Calibration can drift as the data distribution changes.

Evaluate with. Brier score, expected calibration error and reliability curves.

04 · Data design

Signals entering the system

Every useful model depends on the quality, timing and provenance of its inputs. CortexLab treats feature definitions and leakage checks as part of model engineering, not preprocessing trivia.

  • Curated Cortex knowledge — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Projects and innovations — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Published CortexLab articles — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Human-approved corrections — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Conversation feedback — captured with validation, lineage and monitoring so training and production meaning stay aligned.
05 · Evaluation

What has to be measured before calling it successful

  • Retrieval precision
  • Unknown-question rejection
  • Answer helpfulness reward
  • Correction reuse rate
  • Intent accuracy
  • Latency

Evaluation is split between offline quality, operational performance and failure analysis. A high headline metric does not override poor calibration, unstable segments, leakage or unusable latency.

06 · Failure analysis

Where this project can fail

Risk 1

Confidence is as important as retrieval

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 2

Feedback should improve policy separately from facts

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 3

Human corrections need provenance

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

07 · Roadmap

How the project grows without becoming untestable

  1. 01Add trainable intent classifier
  2. 02Train local embeddings
  3. 03Add reranker
  4. 04Train small transformer
  5. 05Distill domain models
  6. 06Offline policy evaluation
08 · Interview lens

Questions an engineer should be able to answer

  • Why is this architecture preferable to a simpler baseline?
  • Which metric can look good while the product still fails?
  • Where can data leakage enter this pipeline?
  • What changes when traffic, data volume or latency requirements increase 10×?
  • Which part should be rolled back first if production quality drops?