Skip to content

The results ledger

Every results file behind a published claim is rendered across the pages below, generated straight from the data in benchmarks/results/: python3 scripts/render_results.py rewrites them and CI fails on drift. Rows are as-run records under the benchmark protocol: quality division numbers never cite timing, perf division numbers name their timing mode, and superseded files are deleted rather than kept beside their replacements, so what is here is the current evidence, whole.

Perf division

bonsai's CUDA growers hold the fastest slot at every measured row scale (10.3s at 16M rows against XGBoost-GPU's 19.6s) at 6.9GB peak host memory against XGBoost's 22.2GB and CatBoost's 19.4GB. On the narrow airline shape bonsai holds both best AUC and fastest fit from 1M rows up. The 2026-07-30 studies hold every width and aspect ratio, with measured device memory that sizes to the problem: 3.4GB at 16M x 128 at constant 2^31-cell volume against XGBoost's 18.9GB and CatBoost's 90.2GB. Every number is same-pod; identical-model GPUs across the rental fleet measure up to ~25% apart.

page what it holds
Fit at scale Row-scale standings: the re-baseline, the XGBoost 3.3 recheck, and the CPU prefetch round.
Width and shape The wide-data arc: the CPU fill, the CUDA recheck, the cols re-baseline, and the iso-volume shape frontier with measured VRAM.
The accuracy-time frontier Accuracy versus fit time at 16M rows, plus the ordered-boosting door.
Airline delays The benchm-ml real-data speed ladder at 0.1M, 1M, and 10M rows.
The single-card ceiling A 500M x 100 matrix trained end to end on one 80GB card.

Quality division

bonsai leads the 55-task Grinsztajn standings at mean rank 1.44 with 36 outright wins; the one knob that translates ambiguously is bracketed in both directions on the standings page. The campaign smoke is the fast local regression check, and the probe archive records every feature declined by measurement.

page what it holds
Grinsztajn standings The only citable standings: 55 third-party tasks, both knob brackets.
Campaign smoke Ten datasets at matched knobs, the fast local regression check.
Probes Every feature admitted or declined by measurement, with its evidence.

The code division

Self-measurement of the bonsai tree, no comparative claim: code metrics.