Search intelligence · Active R&D

SEO Experiment Platform

A controlled optimization platform for titles, metadata, internal links and content refresh decisions using editorial guardrails and bandit-style experimentation.

CTR baselinesBayesian/Thompson-style allocationContent-decay scoringInternal-link similarityMulti-objective reward
7architecture stages
5models / algorithms
7data signals
4evaluation checks
4next milestones
01 · Problem definition

What this system is actually solving

SEO changes are often made without controlled measurement. This project separates hypothesis, experiment, observation and editorial approval.

The project is designed around an observable outcome and explicit constraints. A production decision is only considered successful when the model or algorithm improves a baseline without creating unacceptable cost, latency, reliability or interpretability problems.

02 · Architecture

How the system is decomposed

  1. 01Search Console ingestion
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  2. 02Page opportunity detector
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  3. 03Candidate generator
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  4. 04Approval workflow
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  5. 05Experiment allocator
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  6. 06Reward calculator
    A separately testable stage with defined inputs, outputs, observability and failure handling.
  7. 07Rollback controls
    A separately testable stage with defined inputs, outputs, observability and failure handling.
03 · Models & algorithms

Why each algorithm exists

CTR baselines

Role. CTR baselines is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Bayesian/Thompson-style allocation

Role. Bayesian/Thompson-style allocation is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Content-decay scoring

Role. Content-decay scoring is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Internal-link similarity

Role. Internal-link similarity is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

Multi-objective reward

Role. Multi-objective reward is included because it addresses a specific measurable part of the system rather than being added as decoration.

Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.

Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.

04 · Data design

Signals entering the system

Every useful model depends on the quality, timing and provenance of its inputs. CortexLab treats feature definitions and leakage checks as part of model engineering, not preprocessing trivia.

  • Queries — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Impressions — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Clicks — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Average position — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Engagement — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Conversions — captured with validation, lineage and monitoring so training and production meaning stay aligned.
  • Page age — captured with validation, lineage and monitoring so training and production meaning stay aligned.
05 · Evaluation

What has to be measured before calling it successful

  • Incremental CTR
  • Search traffic quality
  • Conversion delta
  • Core Web Vitals guardrails

Evaluation is split between offline quality, operational performance and failure analysis. A high headline metric does not override poor calibration, unstable segments, leakage or unusable latency.

06 · Failure analysis

Where this project can fail

Risk 1

Traffic alone is a weak reward

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 2

SEO needs guardrails

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

Risk 3

Small samples create false wins

This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.

07 · Roadmap

How the project grows without becoming untestable

  1. 01Query clustering
  2. 02Semantic content gaps
  3. 03Causal uplift estimation
  4. 04Seasonality correction
08 · Interview lens

Questions an engineer should be able to answer

  • Why is this architecture preferable to a simpler baseline?
  • Which metric can look good while the product still fails?
  • Where can data leakage enter this pipeline?
  • What changes when traffic, data volume or latency requirements increase 10×?
  • Which part should be rolled back first if production quality drops?