What we measured, and where it stopped working
This is the whole result set behind every number on the presale page, including the runs that went against us. Nothing here is a projection of your season; all of it is simulation over historical seasons.
Last run before publication · engine v5 · every figure below is a modelled estimate
Method
The engine was tuned on 2018–20 and then frozen. Nothing after 2020 was used to fit it. Everything reported here is out-of-sample on 2021–25.
Every comparison is paired: same seasons, same players, same draft seats, same seeds. A strategy is run against the identical league the baseline saw, so the difference is the strategy rather than the draw.
Draft results are 1,500 drafts per strategy — 5 seasons × 12 seats × 25 seeds. In-season results are 300 simulated seasons. Confidence intervals are 95% and are reported on every effect, including the ones that cross zero.
Results
A random team. One seat in twelve — arithmetic, not a result, and never part of any comparison below.
Draft with v5, then never touch the roster. CI [11.2, 13.3]. The draft engine alone.
Draft with v5 and work the wire, against a passive league. The distance from the middle rung is the +7.0 below.
The middle rung and the +7.0 come from separate experiments with their own runs, which is why 19.4 − 12.2 reads 7.2 here and 7.0 in Part 5's own paired run (19.4 against 12.5). Three tenths of a point is sampling noise; both are reported rather than the flattering one.
Points of roster quality over baseline, 95% CI [+39, +55], across 1,500 drafts per strategy. Playoff rate 71.5% against 67.8% for the app's original ranking.
Points of title probability, 95% CI [+4.7, +9.2], over 300 simulated seasons — a 55% relative increase, and the largest single effect measured anywhere in this backtest. Larger than the entire draft rewrite.
Points of title probability, CI [−3.1, +0.4] — indistinguishable from zero. When all twelve teams run the same waiver policy the edge is gone. The wire is a lever against a passive league and a tax in a sharp one.
Where it does not beat the field
Player projection. The app does not beat expert consensus at predicting player performance. We tested it properly rather than assuming: a walk-forward projection model, 30 features including prior-year usage, scored 0.618 against the consensus's 0.623. It lost. A 200-analyst consensus is not beatable at ranking players off public box-score history, and any claim to the contrary from this data would be overfitting.
What that testing did produce. Four hypotheses tested against the consensus's residual — expert disagreement, touchdown regression, ADP-versus-ECR divergence among them — are real and consistently signed across five seasons. They improved calibration (mean absolute error 54.6 → 52.8) without reordering the board. Better-calibrated points, not better ranks. The rank-to-points curve, fitted on 2016–25 actuals, is the largest single lift in the projection work: correlation 0.522 to 0.623.
Which is why the edge is allocation. The rankings are a commodity — everyone in your league has them. What is ours is what happens after: the rank-to-points curve, the v5 draft allocation, the auction ceilings and survival curves, the waiver policy. Every insurer buys the same weather data; the edge is the underwriting.
The tournament. A better draft reliably gets you into the playoffs. It cannot reliably win them. A six-team bracket is three consecutive single-elimination games; being the best team in the field is worth perhaps 60% per game, and 0.6³ ≈ 22%.
One season. At 12.2% per season, the chance of at least one title is roughly 48% over five years and 73% over ten. Any per-season figure near 80% would require information that does not exist on draft day.
Known limits of this backtest
Simulated opponents are not your opponents. The rivals in these runs bid and claim by fixed policies; a real room is stranger than that, and the draft-side effect is likely smaller against very sharp opposition and larger against very passive opposition.
Five seasons is five seasons. The intervals above are honest about sampling error and silent about a rule change, a scoring change or an injury year that looks like nothing in the historical record.
Roster quality is a modelled measure of a drafted roster, not money won. It converts to titles only through the tournament described above.