Tune the forest for the decision you actually need to make.
A larger forest is not automatically a better forest. Tune tree count, tree size, feature randomness, sample size, and class handling with an evaluation design that matches deployment.
Hyperparameters as levers
First tuning question
Which parameter mainly creates more independent tree opinions, and which parameters control how complex each opinion can be?
| Parameter | Controls | Typical effect |
|---|---|---|
n_estimators | Number of trees | More stable aggregation; more compute and memory. |
max_features | Features considered per split | Lower often reduces correlation but may weaken trees. |
max_depth | Maximum path length | Shallower trees usually reduce variance and complexity. |
min_samples_leaf | Minimum rows in a leaf | Larger leaves smooth predictions and resist tiny pockets. |
max_samples | Rows drawn per bootstrap | Changes tree diversity and training cost. |
class_weight | Relative class costs | Can give rare classes more influence during splits. |
What should we tune first?
Order question
Would you start with 200 combinations, or establish a simple baseline and inspect learning curves first?
- Fit one reproducible baseline with a held-out validation design.
- Check whether adding trees still improves validation performance.
- Try a small grid for
max_features, depth, and leaf size. - Use cross-validation or an appropriate grouped/time split.
- Keep the test set untouched until the final choice.
Metrics must match the business
Metric question
For a rare-fraud detector, can accuracy alone reveal whether the model is useful?
For classification, consider precision, recall, F1, ROC-AUC, PR-AUC, calibration, and the cost of false positives versus false negatives. For regression, consider MAE, RMSE, \(R^2\), subgroup error, and tail behavior.
For imbalanced classification, class_weight="balanced_subsample" adjusts weights from each bootstrap sample. It changes the split objective; it does not replace threshold selection or careful evaluation.