Decision Trees — Splits, Interpretability & When They Fail
A tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Need a model you can read in a review
Prefer
Pruned CART with CV’d depth / ccp_alpha
Greedy local search plus explicit stopping. The explanation is the path. Measure val error before calling it ‘simple.’
- Mixed types without manual crosses.
- A stump is a leakage smoke test.
- Cheap inference on CPU.
Alternative
Unlimited-depth tree because train accuracy is 100%
Low bias, high variance. Leaf constants do not extrapolate. Tomorrow’s drift silently breaks the story.
- Structure flips under a bootstrap of the rows.
- Deep paths are not calibrated probabilities.
- Correlated features make ‘the’ split arbitrary.
Overview
Interviewers use decision trees to test inductive bias, not library calls. A tree greedily partitions feature space into axis-aligned regions. That is why it is readable — and why a single tree is often the wrong production model.
You should be able to:
- Sketch ID3 / C4.5 vs CART.
- Defend Gini vs entropy in one sentence.
- Say when a single tree is enough, and when ensembles exist because it is not.
CART / ID3 intuition
- ID3 / C4.5 style: pick splits that maximize information gain (entropy reduction) for classification.
- CART: binary trees; Gini impurity or MSE for regression; same greedy local search.
- At each node: enumerate candidate thresholds (continuous) or partitions (categorical); choose the split that most purifies children.
- Recurse until purity, min samples, or max depth — then optionally prune.
Gini vs entropy (classification)
- Gini = 1 minus sum of p_k squared — probability two random draws disagree on class; fast; default in many CART implementations.
- Entropy = minus sum of p_k log p_k — information-theoretic; similar splits in practice.
- Interview answer: both measure impurity; Gini is slightly cheaper; depth and sample size dominate the impurity formula.
Depth, pruning, overfitting
- Deep trees memorize noise — low bias, high variance.
- Pre-pruning: max_depth, min_samples_leaf, min_impurity_decrease.
- Post-pruning: grow full, then cost-complexity prune (
ccp_alphain sklearn). - Cross-validate depth /
ccp_alpha; plot train vs val curves.
Categorical vs continuous and missing values
- Continuous: sort values, evaluate midpoints between unique values (or histogram bins).
- Categorical: CART often uses ordered surrogates or one-vs-rest binary splits; high cardinality explodes candidates.
- Missing: surrogate splits, a separate missing branch, or median/mode impute before split — know what your library does.
When a single tree is enough
- Human-readable rules for ops or legal.
- Tiny datasets where ensembles overfit worse.
- Debugging feature pipelines — a stump reveals leakage fast.
- Soft real-time constraints where one shallow tree beats a forest.
When it fails
- Noisy targets → unstable structure under a bootstrap of the data.
- Additive / linear interactions across many weak features → underfit vs GLM or GBM.
- Extrapolation outside seen ranges — trees predict constant leaves.
- Multicollinear features → arbitrary which correlated feature wins the split.
Pros: interpretable path; mixed types; nonlinearities and interactions without manual crosses; cheap inference.
Cons: high variance; greedy local optima; poor probability calibration when deep; unstable importance under correlation.
Architecture (grow, prune, decide)
Decisions
- 1
1 Train rows
- next2 Greedy best split
- 2
2 Greedy best split
- next3 Recurse to depth
- 3
3 Recurse to depth
- next4 Pre or post prune?
- ?
4 Pre or post prune?
- max depth / min leaf5 Stop early
- ccp_alpha5 Cost-complexity prune
- 5
5 Stop early
- next6 CV hyperparameters
- 6
5 Cost-complexity prune
- next6 CV hyperparameters
- 7
6 CV hyperparameters
- next7 Val OK and explainable?
- ?
7 Val OK and explainable?
- yes8 Deploy plus monitor
- no8 Escalate to RF or GBM
- 9
8 Deploy plus monitor
- 10
8 Escalate to RF or GBM
Lesson map
Decision Trees — Splits, Interpretability & When They Fail
A tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB d["1 Train rows"] s["2 Greedy best split"] r["3 Recurse to depth"] p["4 Pre or post prune?"] d -->|1 Train rows to 2 Greedy best split| s s -->|2 Greedy best split| r r -->|3 Recurse to depth| p
Sandbox: impurity + escalate (Python)
Educational helpers. Not sklearn.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same helpers (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Sketch a depth-3 credit tree: income, utilization, missed payments. Walk one applicant down the path as a conjunction. Where does the leaf fail to extrapolate if income is 10x the training max?
Interview Q&A
Why binary CART splits instead of multiway?
Answer
Binary keeps search tractable and avoids fragmenting samples too fast. Multiway categorical splits can starve child nodes.
How do you explain a prediction from a tree?
Answer
Walk the path: each predicate is a conjunction ending in a leaf vote or mean. That conjunction is the explanation — not a SHAP bar chart.
Tree vs logistic regression for a credit policy?
Answer
Tree if nonlinear thresholds and interactions matter and auditors accept rule lists. Logistic if you need stable coefficients and well-calibrated odds with fewer parameters.
What fails at inference time with deep trees?
Answer
Cache-unfriendly deep paths are still usually fast. The real failure is overfitting and brittle updates when data drifts — leaf constants do not extrapolate.
Gini or entropy — which do you pick?
Answer
Either. Splits are similar. Spend the interview on depth, min leaf, and whether a single tree’s variance is acceptable.
How does a stump catch leakage?
Answer
If a one-split tree already crushes the metric, the winning feature is often a post-outcome field or an ID. Fix the data before an ensemble hides it.
Why do ensembles exist if trees are interpretable?
Answer
One tree’s split structure flips when you resample rows. Averaging (RF) or residual fitting (GBM) trades a readable path for lower error. Depth: forests.
High-cardinality category — what goes wrong?
Answer
Too many candidate partitions; tiny child nodes; impurity looks great on a rare level that will not recur. Native categorical handling or a rare-level bucket beats naive one-hot.
Is feature importance from one tree trustworthy?
When is a tree better than GBM for production?
Answer
Tiny n, a hard explainability constraint, or a debug / policy rule you must print. If you already need 20 depth to fit, you wanted an ensemble.