Gridiron Edge ← Back to the presale
Backtest · v5 · published in full

What we measured, and where it stopped working

This is the whole result set behind every number on the presale page, including the runs that went against us. Nothing here is a projection of your season; all of it is simulation over historical seasons.

Published 12 August 2026 · draft engine v5 tuned 2018–20 and tested 2021–25 · weekly model trained 2019–23 and tested 2024

Method

The engine was tuned on 2018–20 and then frozen. Nothing after 2020 was used to fit it. Everything reported here is out-of-sample on 2021–25.

Every comparison is paired: same seasons, same players, same draft seats, same seeds. A strategy is run against the identical league the baseline saw, so the difference is the strategy rather than the draw.

Draft results are 1,500 drafts per strategy — 5 seasons × 12 seats × 25 seeds. In-season results are 300 simulated seasons. Confidence intervals are 95% and are reported on every effect, including the ones that cross zero.

Results

Title probability per season, 12-team league · the three rungs
8.3%

A random team. One seat in twelve — arithmetic, not a result, and never part of any comparison below.

12.2%

The draft engine alone — retired. Measured on a projection source we no longer ship; on the open data, drafting by our board is indistinguishable from drafting by the market. Kept on the chart struck through.

11.5%

A fair seat plus the waiver engine, against a passive league — the +67.5 points a season below — the nine-season re-measure on the shipping engine, positive in every one of them.

The struck rung is the ladder telling the truth about itself: when the draft claim failed re-testing on the open data, the chart kept the corpse rather than quietly re-drawing. The wire effect was re-tested twice in August 2026: a rebuilt five-season harness reproduced +7.0 to the decimal, and then nine seasons took the title figure back down to +3.1 with an interval that touches zero — while the points (+67.5, nine for nine) and playoff odds (63.7% vs 50.0%) replicated in every window. The chart quotes what the fuller data supports, not what the friendlier window said. Moving a number down when the evidence moves is what these pages are for.

Draft engine · a number we retired
+46.6

Measured on a projection source we no longer ship. Re-run on the open data the product actually uses, drafting by our board is statistically indistinguishable from drafting by the market — so the number is retired here, in the same table that once led with it. The draft engine is sold on its mechanism: a live walk-away ceiling for your seat and your room's money.

Waiver policy vs a passive league · nine seasons
+67.5

Points added per season over 540 paired seat-seasons, nine seasons deep: +67.5, positive in all nine (interval +34.1 to +100.9), with playoff odds of 63.7% against a passive 50.0%. The title conversion: 11.5% vs 8.3% — +3.1 points of probability, and the nine-cluster interval (−1.8 to +8.1) contains zero, said plainly. 2018 is the mechanism in one line: +109 points and a worse title outcome — bracket variance, not a broken lever. The points replicate; the ring is a lottery those points buy tickets to.

The same policy in a league where everyone runs it
−1.3

By construction, exactly zero: in a symmetric league titles sum to one, so the all-twelve average gain is an identity, not a finding — the 2026 re-run made the runner say so instead of printing a CI on it. What it measured instead: every seat gains about +35.5 points a season (interval +25.8 to +45.1) and stays at base-rate title odds, and 60% of leagues crown a different champion. The wire is a lever against a passive league and a treadmill in a sharp one.

The weekly model · start/sit calls

The projection result above is why this section exists. If our numbers and the expert rankings tie — and they do — the weekly question stops being who scores more and becomes which of these two do I start. Two separate things live under that heading, and only one of them has been measured. They are kept apart here on purpose.

Measured: start/sit calls. The unit is a single decision between two players, not a lineup and not a season. A weekly projection model was trained on 2019–23 and tested on 2024 — one held-out season, a different and shorter window than the draft engine's 2018–20 tune and 2021–25 test above. That test season contains 3,902 player-weeks, which produce 133,388 pairwise start/sit decisions. Every figure below is a hit rate over those pairs.

Shipped, and still unmeasured: the win-probability lineup solve. The thing described on the front page is live in the current build: the app reads your actual opponent from the schedule, simulates your floor lineup and your ceiling lineup against that roster at ten thousand runs each, and states the trade-off in its own computed numbers — the format it renders, for illustration, being “CEILING wins 34% vs FLOOR's 31% — you're the underdog.” When the two lineups differ by less than simulation noise it says so instead of manufacturing a preference; when no opponent can be read it refuses out loud and falls back to season-strength advice. Those are per-matchup computations, on demand, not averages of anything.

What we still do not claim is a season-long value for matchup switching. Nobody has measured whether switching by matchup wins you more games across a season, so no aggregate, win count or percentage for it appears anywhere on this site. When the season-long runs exist they will appear here first, whichever way they come out.

Against a recency baseline · last three weeks of production
60.2%

Of 133,388 pairwise start/sit decisions across 3,902 player-weeks in the 2024 test season, 60.2% land on the right side when scored against a board built from each player's last three weeks of production. That is the baseline most managers actually use, which is why it is reported first.

Backtest measurement · out-of-sample 2024 · N = 133,388 pairs

Against season-average rankings · the disagreement subset
51.0%

A season-average board is the harder baseline, and against it we mostly agree: the two boards pick the same starter on over 97% of pairs. On the contested pairs, the first test season read 53.2% — and the caveat under it said one season could not separate that from a coin flip. Four more seasons settled it: 51.0% across five folds. The caveat was right, the claim is retired, and this row exists so you can watch that happen.

The frequency matters as much as the rate. At 2.4% of pairs, a typical roster meets a contested call a handful of times a season — three to five, not weekly. Anyone buying this for one week should read that as a warning rather than a feature.

Backtest measurement · out-of-sample 2024 · N = 3,201 contested pairs

These are hit rates, not margins. A right call in a week you lose by forty still counts as right, because the decision was the only thing in our control. Nothing here says how much a right call is worth in points.

One test season is one test season. The draft results above have five held-out seasons behind them, and as of August 2026 so does this one — which is how the 53.2% became a 51.0% and got retired. What survives the same five folds: contested calls against the trailing-three-week board most managers actually run land 60.2%, above 50% in every season including one the model had never seen.

Late news is not isolated here. The re-solve on final inactives is part of the product but is not a measured effect in these runs; the figures use the information available at the historical lock time.

Where it does not beat the field

Player projection. The app does not beat expert rankings at predicting player performance. We tested it properly rather than assuming: a walk-forward projection model — forty-one features tried, including snap shares, draft capital, red-zone role and Vegas lines — with the shipping open model scoring 0.583 against the expert consensus' 0.623. It lost, and the expert-grade features added nothing the market had not already priced. Both results are in the run log. Ranking players off public box-score history is a crowded, well-solved problem, and any claim to the contrary from this data would be overfitting.

What that testing did produce. Four hypotheses tested against the consensus's residual — expert disagreement, touchdown regression, ADP-versus-ECR divergence among them — are real and consistently signed across five seasons. They improved calibration (mean absolute error 54.6 → 52.8) without reordering the board. Better-calibrated points, not better ranks. The rank-to-points curve, fitted on 2016–25 actuals, is the largest single lift in the projection work: correlation 0.522 to 0.623.

Which is why the edge is the decision. The rankings are a commodity — everyone in your league has them. What is ours is what happens after: the rank-to-points curve, the weekly win-probability solve above, the v5 draft allocation, the auction ceilings and survival curves, the waiver policy.

The tournament. A better draft reliably gets you into the playoffs. It cannot reliably win them. A six-team bracket is three consecutive single-elimination games; being the best team in the field is worth perhaps 60% per game, and 0.6³ ≈ 22%.

One season. At 11.5% per season — the nine-season rung — the chance of at least one title is roughly 46% over five years and 71% over ten. Any per-season figure near 80% would require information that does not exist on draft day.

Known limits of this backtest

Simulated opponents are not your opponents. The rivals in these runs bid and claim by fixed policies; a real room is stranger than that, and the draft-side effect is likely smaller against very sharp opposition and larger against very passive opposition.

Five seasons is five seasons. The intervals above are honest about sampling error and silent about a rule change, a scoring change or an injury year that looks like nothing in the historical record.

Roster quality is a modelled measure of a drafted roster, not money won. It converts to titles only through the tournament described above.

See pricing What this will not do for you
Not buying today?

One email before your draft, one before the price moves. Nothing else, and one click to stop.

{{ signupNote }}

Gridiron Edge is operated by Dave Mendlen, Seattle, Washington Terms & refund policy Privacy support@gridironedge.app