Choosing ML Algorithms
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
Choosing ML Algorithms
6 studies- 1.Choosing ML Algorithms — Problem Shape to Model FamilySenior interviewers do not want an algorithm encyclopedia. They want you to map problem shape to model family under constraints: tabular vs sequence, interpretability vs accuracy, latency, data size, and drift. Start with a baseline; ship the simplest model that meets the metric and the ops budget.
- 2.Decision Trees — Splits, Interpretability & When They FailA tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
- 3.Random Forests & Bagging — Variance Reduction & Feature ImportanceA senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
- 4.Gradient Boosted Trees — XGBoost / LightGBM / CatBoost Tradeoffs“Just use XGBoost” is not a senior answer. Explain sequential residual fitting, the learning-rate × n_estimators tradeoff, early stopping, and honest differences among XGBoost, LightGBM, and CatBoost — then defend when GBM beats RF, and when calibration or latency make RF or linear wiser.
- 5.Time Series Forecasting — Classical, ML Features & Deep SequencesTell a truly temporal problem from a table that happens to have a timestamp. Wrong choice means leakage, optimistic MAPE, and models that die after a holiday. Map classical ARIMA/ETS/Prophet, trees on lags, and deep sequences — with walk-forward validation. Prediction is not a counterfactual.
- 6.Algorithm Selection Playbook — Metrics, Baselines & Failure ModesChoosing XGBoost is easy; defending the full selection loop is not. Start with baselines, pick metrics that match business cost, handle imbalance, run a leakage checklist, know when not to use deep learning, and design for retrain cadence, drift, and compliance explainability.