Algorithm Selection Playbook — Metrics, Baselines & Failure Modes
Choosing XGBoost is easy; defending the full selection loop is not. Start with baselines, pick metrics that match business cost, handle imbalance, run a leakage checklist, know when not to use deep learning, and design for retrain cadence, drift, and compliance explainability.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you run the bake-off
Prefer
Cost-matched metric, leakage-safe split, then linear → tree → RF → GBM / TS
If the fancy model cannot beat a dumb baseline on a clean validation design, stop and fix data. Shadow deploy the champion.
- Accuracy is a trap under imbalance.
- Thresholds belong on validation, not folklore 0.5.
- Drift monitors and a warm previous champion are launch requirements.
Alternative
Train GBM once, optimize accuracy, skip the baseline
False wins. Silent leakage. Deep learning on 8k tabular rows. SHAP as a causal story for compliance.
- No retrain trigger, no rollback.
- Random splits on users or time.
- Complexity without a business delta.
Selection loop
Negotiate the metric early. Complexity is a cost.
- 1
Frame cost
False negative dollars ≠ false positive dollars. Write the operating point. - 2
Baseline
Mean / majority / linear / seasonal naive. Log it in the tracker. - 3
Split fairly
Walk-forward or group by user / session. No random K-fold on time or entities. - 4
Escalate families
Linear, then tree / RF, then GBM or a TS family from the earlier lessons. - 5
Audit and ship
Leakage, calibration, slices. Shadow, monitor drift, keep a rollback.
Overview
Choosing XGBoost is easy; defending the full selection loop in a senior interview is not. This playbook starts with baselines, picks metrics that match business cost, handles imbalance, runs a leakage checklist, and designs for retrain cadence, drift, and compliance.
Always start with a baseline
- Mean / median predictor (regression); majority / stratified random (classification).
- Linear / logistic with simple features.
- Seasonal naive for time series.
- If your fancy model cannot beat these convincingly on a clean validation design, stop and fix data.
Metric choice (defend it)
- Ranking / discrimination: AUC-ROC, AUC-PR (PR when positives are rare).
- Thresholded decisions: F1, precision at k, recall at fixed precision — match the operating point.
- Regression: MAE (robust), RMSE (penalizes large errors), MAPE (careful with zeros / scale).
- Business cost: false negative $ ≠ false positive $ — optimize expected cost or use cost-sensitive learning.
- Calibration (Brier, ECE) when probabilities drive pricing or capacity.
Class imbalance
- Accuracy is a trap; report minority recall / precision and PR curves.
- Resampling (under / over), class weights, or focal-style losses — try simplest weight first.
- Threshold tuning on validation beats chasing rare tricks.
Leakage checklist
- Target encoding / scalers fit on full data including test.
- Random splits on time-ordered or grouped entities (users, sessions).
- Features that are consequences of the label (post-outcome fields).
- Duplicate or near-duplicate rows across train / test.
- Future covariates treated as known (see forecasting).
When NOT to use deep learning
- Flat tabular, modest n, need audits — prefer linear / trees / GBM.
- Latency / energy budgets on CPU edge.
- Label noise dominates; more capacity memorizes noise.
- Team cannot operate GPUs, monitoring, and continual retraining.
- A linear model already meets the metric — complexity is unjustified.
Production constraints
- Retrain cadence: batch daily vs online learning; feature store as-of correctness.
- Concept drift: monitor input distributions + performance; triggers for retrain.
- Explainability for compliance: coefficients, sparse trees, SHAP with caveats — document limits.
- Shadow deploy / A/B; keep the previous champion warm for rollback.
Pros of a disciplined playbook: fewer false wins; clearer stakeholder stories; safer launches.
Cons: feels slower than “train GBM once”; requires metric negotiation early.
Architecture (frame, fit fairly, ship)
Decisions
- 1
1 Business cost + operating point
- next2 Choose metric + baseline
- 2
2 Choose metric + baseline
- next3 Leakage-safe split
- 3
3 Leakage-safe split
- next4 Linear then tree then GBM
- 4
4 Linear then tree then GBM
- next5 Leakage + calibration + slices
- 5
5 Leakage + calibration + slices
- next6 Beat agreed baseline?
- ?
6 Beat agreed baseline?
- no7 Fix data / features
- yes7 Shadow + drift + retrain
- 7
7 Fix data / features
- next3 Leakage-safe split
- 8
7 Shadow + drift + retrain
- next8 Explainability if required
- 9
8 Explainability if required
Lesson map
Algorithm Selection Playbook — Metrics, Baselines & Failure Modes
Choosing XGBoost is easy; defending the full selection loop is not. Start with baselines, pick metrics that match business cost, handle imbalance, run a leakage checklist, know when not to use deep learning, and design for retrain cadence, drift, and compliance explainability.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB biz["1 Business cost + operating point"] met["2 Choose metric + baseline"] split["3 Leakage-safe split"] fit["4 Linear then tree then GBM"] biz -->|1 Business cost + operating point| met met -->|2 Choose metric + baseline| split split -->|3 Leakage-safe split| fit
Sandbox: majority baseline + cost (Python)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same helpers (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
90% of users never churn. A PM wants “highest accuracy.” Compute the majority baseline. Propose a metric and a threshold plan for a retention discount that costs $20 and saves $80 if it works — without claiming the churn score is causal.
Interview Q&A
Stakeholder wants ‘highest accuracy.’ What do you do?
Answer
Translate to operating costs and class priors; propose AUC-PR or cost curves; show majority baseline accuracy to reset expectations.
How do you detect leakage quickly?
Answer
Suspiciously high metrics; features with impossible as-of times; identical IDs across splits; ablation of “too good” features collapses performance.
When is RF enough and GBM overkill?
What belongs in a model card / handoff?
Answer
Metric definition, split scheme, baselines beaten, known failure slices, retrain triggers, and explainability method with limitations.
Accuracy is 95% on 5% positives. Celebrate?
Answer
No. Majority baseline is already ~95%. Look at PR, minority recall, and dollars. Tune the threshold on validation.
User-level data, random row split?
Answer
Rows from the same user leak. Group / block by entity. Time-ordered outcomes also need walk-forward. Depth: forecasting.
Need probabilities for capacity planning?
Answer
Then calibration is part of the metric. Brier / ECE plus reliability plots. A high AUC model can still be overconfident.
Compliance wants ‘the important features.’
Answer
Coefficients or a pruned tree if the function is simple; permutation / SHAP with correlation caveats otherwise. Do not tell a causal story from a predictive rank. Trees if you actually need rules.
When do you refuse deep learning?
Answer
Modest flat tabular, CPU edge, noisy labels, no GPU ops story, or a linear model already inside the delta. Vision / NLP / long multimodal sequences are where capacity earns its keep. Hub: shape → family.
What is a retrain trigger?
Answer
PSI / population shift on key inputs, a drop on a delayed label slice, or a calendar (weekly job). Keep the previous champion warm. No trigger is how silent drift ships.