What this system is actually solving
Businesses often know a customer churned only after the relationship is lost. This project frames churn as a time-aware prediction problem with intervention cost, class imbalance and explainability requirements.
The project is designed around an observable outcome and explicit constraints. A production decision is only considered successful when the model or algorithm improves a baseline without creating unacceptable cost, latency, reliability or interpretability problems.
How the system is decomposed
- 01Event ingestion and customer feature store
A separately testable stage with defined inputs, outputs, observability and failure handling. - 02Time-aware train/validation split
A separately testable stage with defined inputs, outputs, observability and failure handling. - 03Baseline logistic regression
A separately testable stage with defined inputs, outputs, observability and failure handling. - 04Tree ensemble candidate models
A separately testable stage with defined inputs, outputs, observability and failure handling. - 05Probability calibration
A separately testable stage with defined inputs, outputs, observability and failure handling. - 06SHAP-style explainability layer
A separately testable stage with defined inputs, outputs, observability and failure handling. - 07Risk-segment API and monitoring dashboard
A separately testable stage with defined inputs, outputs, observability and failure handling.
Why each algorithm exists
Logistic regression
Role. Logistic regression is included because it addresses a specific measurable part of the system rather than being added as decoration.
Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.
Random forest
Role. Random forest is included because it addresses a specific measurable part of the system rather than being added as decoration.
Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.
Gradient boosting
Role. Gradient boosting is included because it addresses a specific measurable part of the system rather than being added as decoration.
Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.
Calibration curves
Role. Aligns predicted confidence with observed outcome frequency so thresholds mean something operationally.
Trade-off. Calibration can drift as the data distribution changes.
Evaluate with. Brier score, expected calibration error and reliability curves.
Cost-sensitive thresholding
Role. Cost-sensitive thresholding is included because it addresses a specific measurable part of the system rather than being added as decoration.
Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.
Population stability monitoring
Role. Population stability monitoring is included because it addresses a specific measurable part of the system rather than being added as decoration.
Trade-off. The component must be compared with a simpler baseline and removed if it adds complexity without measurable value.
Evaluate with. Task-specific quality metric, latency, reliability and failure-case analysis.
Signals entering the system
Every useful model depends on the quality, timing and provenance of its inputs. CortexLab treats feature definitions and leakage checks as part of model engineering, not preprocessing trivia.
- Account tenure — captured with validation, lineage and monitoring so training and production meaning stay aligned.
- Billing and plan changes — captured with validation, lineage and monitoring so training and production meaning stay aligned.
- Product usage frequency — captured with validation, lineage and monitoring so training and production meaning stay aligned.
- Support interactions — captured with validation, lineage and monitoring so training and production meaning stay aligned.
- Recent engagement change — captured with validation, lineage and monitoring so training and production meaning stay aligned.
- Historical churn labels — captured with validation, lineage and monitoring so training and production meaning stay aligned.
What has to be measured before calling it successful
- PR-AUC for imbalanced performance
- Recall at operational capacity
- Expected retention value
- Calibration error
- Segment fairness checks
Evaluation is split between offline quality, operational performance and failure analysis. A high headline metric does not override poor calibration, unstable segments, leakage or unusable latency.
Where this project can fail
Why accuracy can mislead on churn
This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.
Why thresholds are business decisions
This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.
How leakage occurs in customer features
This risk is tracked through tests, monitoring, explicit thresholds or human review depending on where it appears in the architecture.
How the project grows without becoming untestable
- 01Add survival analysis
- 02Introduce uplift modeling
- 03Build intervention recommender
- 04Monitor drift and retrain triggers
Questions an engineer should be able to answer
- Why is this architecture preferable to a simpler baseline?
- Which metric can look good while the product still fails?
- Where can data leakage enter this pipeline?
- What changes when traffic, data volume or latency requirements increase 10×?
- Which part should be rolled back first if production quality drops?