6. Regularization & constraints¶
The idea¶
Boosting will happily memorize the training set; every knob here limits how it's allowed to fit. Three families:
- Shrink the leaves: L2 (
lambda_l2) and L1 (lambda_l1) penalties on leaf values. - Starve the splits:
min_child_hess,min_data_in_leaf,min_gain_to_split,max_depth/max_leaves, and per-tree feature subsampling (feature_fraction). - Constrain the shape: monotone constraints (predictions must be non-decreasing/-increasing in a feature) and interaction constraints (features may only combine within declared groups). These encode domain knowledge rather than fight variance.
The math¶
L2 appears as the \(+\lambda\) in every score and leaf value: it damps leaves supported by little evidence. L1 soft-thresholds the gradient sum:
so leaves with \(|G| \le \alpha\) are exactly zero (sparsity), and everything else shrinks toward zero by \(\alpha\).
Monotone (\(+1\) on feature \(j\)): every split on \(j\) must have \(w_L \le w_R\), and that must keep holding as descendants refine, so after a constrained split, the left subtree's values are capped and the right subtree's floored at the midpoint \((w_L + w_R)/2\). Bounds inherit downward; leaf values clamp into them.
Interaction groups: a node may split on feature \(f\) only if some declared group contains \(f\) together with every feature already used on the path from the root (features outside all groups may keep splitting alone). Prediction becomes a sum of functions over the groups only.
In bonsai¶
- L1/L2:
l1_thresholded, the four-argumentscore, andbounded_leaf_weightininclude/bonsai/split.hpp. The same three functions serve the finders and the growers, so gain and leaf values can't disagree about the penalty. - feature_fraction:
sample_featuresinsrc/grower.cpp: a per-tree sorted draw from a grower-owned rng (tree.feature_seed); unselected features get zero-binned placeholder histograms the finders skip. - Monotone: two touch points. Rejection: in the candidate loop of
src/split.cpp, skip when \(\text{mc} \cdot (w_R - w_L) < 0\), using bounded child weights. Propagation:propagate_monotone_boundsinsrc/grower.cppfences children at the midpoint viaSplitInput::lo/hi;finalize_as_leafclamps into them. Config:tree.monotone_constraints = [1, 0, -1, ...](or--set tree.monotone_constraints=1,0,-1). - Interaction:
SplitInput::allowed/pathcarry the permitted set down the tree;allowed_features/propagate_interaction_state(src/grower.cpp) recompute it per split; the finder masks excluded features. Config:tree.interaction_constraints = ["0,1", "2,3"](CLI:0+1,2+3).
The oblivious grower rejects both constraint types at construction rather than silently ignoring them.
Try it¶
import numpy as np
import bonsai
rng = np.random.default_rng(0)
X = rng.normal(size=(4000, 8)).astype(np.float32)
y = (X[:, 0] * 2.0 + X[:, 1] + rng.normal(0, 0.1, 4000)).astype(np.float32)
# Monotone: prediction non-decreasing in feature 0.
mono = bonsai.BonsaiRegressor(
n_iters=60, grower="depthwise",
params={"tree.monotone_constraints": "1,0,0,0,0,0,0,0"}).fit(X, y)
# Interaction: features 0-3 may not mix with 4-7 on any path.
inter = bonsai.BonsaiRegressor(
n_iters=60, grower="depthwise",
params={"tree.interaction_constraints": "0+1+2+3,4+5+6+7"}).fit(X, y)
print("monotone R2:", round(mono.score(X, y), 4))
print("interaction R2:", round(inter.score(X, y), 4))
The unit tests are readable specs: [grower][monotone] asserts a
non-monotone dataset yields a provably monotone prediction curve;
[grower][interaction] walks every root-to-leaf path and asserts groups
never mix (tests/unit/test_grower.cpp).
Gotchas & war stories¶
- Rejecting the split isn't enough for monotonicity. A split on an
unconstrained feature can still create descendants that later violate
the constrained feature's ordering; that's what the inherited
lo/hibounds prevent. Basic-mode implementations that skip propagation produce trees that are monotone split-by-split and non-monotone end-to-end. - L1's scale is data-scale.
lambda_l1=100was needed on YearPredictionMSD (leaf gradient sums in the thousands) to move RMSE at all; the same value on California Housing would zero half the leaves. - A feature outside every interaction group isn't banned: it can still split, just never with anything else on the path.
Appendix: why the midpoint fence is sufficient¶
The claim to prove: with constraint \(+1\) on feature \(j\), the whole tree's prediction is non-decreasing in \(x_j\), not merely each individual split.
Take two inputs \(x, x'\) identical except \(x_j \le x'_j\). Splits on any other feature route them identically, so their root-to-leaf paths can only diverge at splits on \(j\), and there, \(x\) goes left whenever \(x'\) does not go further left, i.e. at the topmost divergence \(x\) enters the left child and \(x'\) the right.
Now the fence does its work. At that constrained split with (bounded) child weights \(w_L \le w_R\), the left subtree inherits the ceiling \(m = (w_L + w_R)/2\) and the right subtree the floor \(m\); every descendant leaf clamps into its inherited interval, and deeper constrained splits only narrow it. So \(x\)'s leaf value is \(\le m\) and \(x'\)'s is \(\ge m\), giving \(f(x) \le f(x')\) for any pair, hence monotonicity end-to-end. \(\blacksquare\)
Why the midpoint rather than fencing left at \(w_L\) and right at \(w_R\)? Any value in \([w_L, w_R]\) makes the proof go through; the midpoint just splits the remaining headroom evenly so neither subtree starts its refinement already pinned to its bound. And the proof shows exactly why rejection-only implementations fail (gotchas): without inherited bounds, a later, unconstrained split can push a left-subtree leaf above \(m\): each split individually fine, the composition broken.