Random Forests & Bagging — Variance Reduction & Feature Importance
A senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why not one deep tree?
Prefer
Bagged, feature-randomized trees (RF)
High-variance bases, averaged. Feature bagging stops every tree from cloning the same strong split.
- Embarrassingly parallel train.
- OOB error as an early thermometer.
- Fewer knobs than a well-tuned GBM.
Alternative
One unpruned tree, or RF importances as a causal story
A single tree’s structure flips under resample. MDI splits importance across correlated blocks and favors high-cardinality features.
- OOB is not a test set.
- 500 deep trees can miss a tight p99.
- Tabular SOTA often still belongs to GBM.
Overview
A senior interview answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.”
You should be able to:
- Sketch bootstrap aggregating and feature bagging.
- Say what OOB is and is not.
- Report importance without pretending it is causal.
Bootstrap aggregating (bagging)
- Draw B bootstrap samples (sample n rows with replacement).
- Fit a base learner (usually an unpruned or deep tree) on each.
- Aggregate: majority vote (classification) or mean (regression).
- Bagging helps high-variance, low-bias bases; little help for already-stable models.
Feature bagging (the “random” in RF)
- At each split, consider only mtry random features (classification often sqrt(p); regression often p/3 — tune).
- Decorrelates trees so averages reduce variance more than identical bagged trees.
- Extra-Trees / extremely randomized trees push randomness further onto thresholds.
OOB error
- Each tree’s unused ~37% samples form a free validation stream.
- Aggregate OOB predictions → rough generalization estimate without a holdout.
- Useful early signal; still confirm with a proper validation set (especially under leakage risk).
Feature importance pitfalls
- Mean decrease impurity (MDI): biased toward high-cardinality / continuous features.
- Correlated features: importance splits across the group — none looks “important.”
- Prefer permutation importance on a validation set; still correlational, not causal.
- SHAP / group permutation for correlated blocks when stakeholders demand ranks.
When RF beats a single tree
- Noisy labels or unstable splits.
- Need a strong default with little hyperparameter drama.
- Parallel training across trees; embarrassingly parallel.
- Probability ranks more stable than one deep tree (still calibrate if thresholds matter).
When RF loses to boosting
- Many tabular competitions and industry tabular SOTA favor GBM.
- RF underfits complex additive structure that sequential residual fitting captures.
- Large sparse high-dimensional text: linear models or boosted linear often win; RF struggles with very high p without care.
Parallelism and inference cost
- Train: near-linear speedup with cores until memory bandwidth saturates.
- Infer: sum/vote over hundreds of deep trees — can miss tight p99 budgets vs a shallow GBM or distilled model.
- Mitigation: fewer / shallower trees, early exit, or replace with LightGBM for latency.
Pros: robust default; OOB; nonlinearities; parallel; fewer catastrophic overfits than one deep tree.
Cons: less interpretable than one tree; weaker than well-tuned GBM on many tabular tasks; heavy inference; MDI importance misleading.
Architecture (bag, diagnose, ship)
Decisions
- 1
1 Train matrix
- next2 Bootstrap B samples
- 2
2 Bootstrap B samples
- next3 Grow deep trees with mtry
- 3
3 Grow deep trees with mtry
- next4 OOB error
- next4 Permutation importance
- 4
4 OOB error
- next6 Beat linear and tree?
- 5
4 Permutation importance
- next5 Flag correlated groups
- 6
5 Flag correlated groups
- next6 Beat linear and tree?
- ?
6 Beat linear and tree?
- yes and latency OK7 Deploy RF plus monitor
- accuracy gap7 Try gradient boosting
- need rules7 Shallow surrogate tree
- 8
7 Deploy RF plus monitor
- 9
7 Try gradient boosting
- 10
7 Shallow surrogate tree
Lesson map
Random Forests & Bagging — Variance Reduction & Feature Importance
A senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB x["1 Train matrix"] b["2 Bootstrap B samples"] t["3 Grow deep trees with mtry"] oob["4 OOB error"] x -->|1 Train matrix to 2 Bootstrap B samples| b b -->|2 Bootstrap B samples| t t -->|3 Grow deep trees with mtry| oob
Sandbox: OOB fraction + picker (Python)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same helpers (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Two features are 0.95 correlated and both predictive. Sketch why bagging rows alone still yields similar trees, then what mtry changes. How would you report importance to a PM without sounding causal?
Interview Q&A
Why does feature bagging matter beyond row bagging?
Answer
Row bagging alone yields correlated trees if the same strong features dominate every split. Restricting features per split decorrelates trees so averaging reduces variance more.
Is OOB a replacement for a test set?
Answer
No — convenient estimate during training. Still hold out a true test, and never use OOB if preprocessing leaked across the full matrix.
Why can RF look worse than a tuned GBM?
Answer
Boosting reduces bias by fitting residuals sequentially; RF mostly averages variance. Complex additive structure often favors GBM. Depth: boosting.
How do you report feature importance safely?
Answer
Prefer permutation importance on validation; group correlated features; state non-causality; for compliance pair with domain review. Treatment questions go to Causal ML, not MDI.
What is Extra-Trees adding?
Answer
Randomized thresholds on top of random feature subsets. More bias, often less variance, faster train. Still not a substitute for a leakage-safe split.
Classification mtry heuristic?
Answer
sqrt(p) is a starting point, not a law. Tune when many weak features or a few dominant ones. Measure OOB and a holdout.
RF for text with 100k sparse tokens?
Answer
Usually no. Linear models or linear boosting on sparse features beat RF, which struggles when p is huge and signal is additive in bag-of-words space.
Need rules but RF won on val?
Answer
Extract a shallow surrogate tree or a rule list on RF predictions. Do not pretend 300 trees are a policy document. Depth: trees.
Does more trees always help?
Answer
Error typically plateaus; extra trees cost RAM and p99. Plot OOB vs n_estimators and stop when the curve flattens — unlike GBM, extra RF trees rarely overfit the same way, but they are not free.
When is RF the model you ship under time pressure?
Answer
Small team, CPU parallelism, limited tuning window, and RF already inside the business delta of a GBM. Monitor; challenge with boosting later. Depth: playbook.