Model Evolution
The record of what this model picks, why, what happened, and what changes next — including everything that didn't work. Every number here is measured walk-forward: each season predicted by a fit that never saw it. A model scored on data it trained on can be made to look like anything.
Current Model
market+injuries/v1 · active since 2026-09-14
Reads: de-vigged market moneyline (logit), injury out differential, injury questionable differential
First model to use the betting line at all. The site previously ranked teams with Elo alone, which measured 62.52% -- so this is a +4.5 point gain over our own baseline, not over the market. Injuries are included because they are real information that costs nothing to add, not because they were shown to help: their fitted weights are near zero (out -0.004, questionable -0.022) against market_logit at 0.954.
67.00% straight-up · log loss 0.60687 · 72.6% of Pick'em points · 3,130 held-out games
The Honest Position
This model does not beat the market. Across 3,348 held-out games it picked 66.67% against the market's 66.58% — a difference of +0.09 points, which is noise.
The deeper problem is disclosed rather than hidden: on the 53 disagreements the model was right 52.8% of the time (z=0.41); within what noise explains
A model that almost never disagrees with the line cannot beat it — it is closer to a market tracker than an independent opinion. Fixing that is the first item on the improvement list, and it is the reason the gain worth claiming is over our own previous Elo baseline (62.52%), not over the market.
Weekly Edge Search
Every week the model is tested against the market across real slices of the schedule. Comparing raw accuracy is close to useless, because on almost every game both pick the same side. The question that carries information is: on the games where they disagree, who is right? With no edge, the answer is 50%. A slice is only called an edge with at least 30 disagreements and a z-score beyond 2 — because a 3-from-4 record is exactly what noise looks like.
Nothing beat the market beyond noise in the latest run. That is the result, recorded rather than reframed. Fourteen slices were tested; finding one "significant" result among fourteen is what multiple comparisons produce on their own, which is why a candidate has to repeat on fresh data before it counts.
| Segment | Games | Model | Market | Edge | What the disagreements showed |
|---|---|---|---|---|---|
| all games | 3348 | 66.7% | 66.6% | +0.09 | on the 53 disagreements the model was right 52.8% of the time (z=0.41); within what noise explains |
| home underdogs | 2104 | 67.3% | 67.3% | +0.00 | on the 32 disagreements the model was right 50.0% of the time (z=0.00); within what noise explains |
| road favourites | 2104 | 67.3% | 67.3% | +0.00 | on the 32 disagreements the model was right 50.0% of the time (z=0.00); within what noise explains |
| near pick-ems (3 or less) | 1259 | 55.4% | 55.1% | +0.24 | on the 53 disagreements the model was right 52.8% of the time (z=0.41); within what noise explains |
| division rivalries | 1193 | 67.2% | 67.1% | +0.17 | only 16 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| off a bye | 992 | 66.0% | 65.2% | +0.81 | only 18 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| late season (weeks 14+) | 991 | 69.6% | 69.2% | +0.40 | only 10 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| big spreads (7+) | 966 | 80.5% | 80.5% | +0.00 | only 0 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| indoors | 950 | 66.5% | 66.7% | -0.21 | only 14 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| long travel (2000km+) | 838 | 67.8% | 67.7% | +0.12 | only 19 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| heavy injury gap (3+) | 786 | 67.8% | 67.9% | -0.13 | only 11 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| early season (weeks 1-4) | 778 | 62.7% | 62.5% | +0.26 | only 14 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| a quarterback is out | 478 | 70.5% | 71.3% | -0.84 | only 12 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
| short week | 210 | 71.9% | 72.4% | -0.48 | only 1 games where the model and the market disagreed -- below the 30 needed to tell an edge from noise |
Grading Every Pick
11/15 correct (73.3%)
With the market: 0 right, 0 wrong. Against the market: 11 right, 4 wrong.
The second line is the one that matters. Losing a game the whole market also picked is an upset, not a modelling error — no reweighting would have caught it, and tuning on those is how a model gets fitted to noise. On the 15 games where we actually took a different view, we were right 73.3% of the time.
2026 Week 1 · 0 pts · Our error
Picked LAC at 52.6% — lost, ARI won.
Why we picked it: The market made LAC a 0.0% favourite, and the injury reports moved it up 52.6 points (2 more away players ruled out), for a final 52.6%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We disagreed with the market and lost. This is our error, not an upset. The injury adjustment pushed us 52.6 points off the market's number and it was the wrong direction -- exactly the case for weighting injuries less, not more.
2026 Week 1 · 0 pts · Right, against the market
Picked PIT at 56.0% — correct.
Why we picked it: The market made PIT a 0.0% favourite, and the injury reports moved it up 56.0 points (4 more away players ruled out), for a final 56.0%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked BAL at 56.4% — correct.
Why we picked it: The market made BAL a 0.0% favourite, and the injury reports moved it up 56.4 points (1 more away players ruled out; 1 more questionable on the away side), for a final 56.4%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked BUF at 54.0% — correct.
Why we picked it: The market made BUF a 0.0% favourite, and the injury reports moved it up 54.0 points (3 more questionable on the away side), for a final 54.0%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked CHI at 56.3% — correct.
Why we picked it: The market made CHI a 0.0% favourite, and the injury reports moved it up 56.3 points (1 more home players ruled out; 4 more questionable on the away side), for a final 56.3%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked JAX at 57.0% — correct.
Why we picked it: The market made JAX a 0.0% favourite, and the injury reports moved it up 57.0 points (1 more away players ruled out), for a final 57.0%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked NYG at 54.1% — correct.
Why we picked it: The market made NYG a 0.0% favourite, and the injury reports moved it up 54.1 points (1 more away players ruled out; 1 more questionable on the home side), for a final 54.1%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked MIN at 56.9% — correct.
Why we picked it: The market made MIN a 0.0% favourite, and the injury reports moved it up 56.9 points (1 more away players ruled out; 3 more questionable on the away side), for a final 56.9%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Our error
Picked MIA at 59.6% — lost, LV won.
Why we picked it: The market made MIA a 0.0% favourite, and the injury reports moved it up 59.6 points (1 more home players ruled out), for a final 59.6%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We disagreed with the market and lost. This is our error, not an upset. The injury adjustment pushed us 59.6 points off the market's number and it was the wrong direction -- exactly the case for weighting injuries less, not more.
2026 Week 1 · 0 pts · Right, against the market
Picked SEA at 52.5% — correct.
Why we picked it: The market made SEA a 0.0% favourite, and the injury reports moved it up 52.5 points (1 more away players ruled out; 2 more questionable on the home side), for a final 52.5%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked DET at 51.7% — correct.
Why we picked it: The market made DET a 0.0% favourite, and the injury reports moved it up 51.7 points (3 more away players ruled out), for a final 51.7%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Our error
Picked TEN at 62.1% — lost, NYJ won.
Why we picked it: The market made TEN a 0.0% favourite, and the injury reports moved it up 62.1 points (3 more away players ruled out; 2 more questionable on the away side), for a final 62.1%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We disagreed with the market and lost. This is our error, not an upset. The injury adjustment pushed us 62.1 points off the market's number and it was the wrong direction -- exactly the case for weighting injuries less, not more.
2026 Week 1 · 0 pts · Right, against the market
Picked SF at 53.9% — correct.
Why we picked it: The market made SF a 0.0% favourite, and the injury reports moved it up 53.9 points (1 more questionable on the away side), for a final 53.9%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Right, against the market
Picked CIN at 53.8% — correct.
Why we picked it: The market made CIN a 0.0% favourite, and the injury reports moved it up 53.8 points (2 more away players ruled out; 3 more questionable on the away side), for a final 53.8%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We took a different side from the market and were right. This is the ONLY kind of result that is evidence of an edge. Worth checking whether the same signal recurs -- one of these is luck, a pattern of them is an edge.
2026 Week 1 · 0 pts · Our error
Picked WAS at 52.0% — lost, PHI won.
Why we picked it: The market made WAS a 0.0% favourite, and the injury reports moved it up 52.0 points (2 more home players ruled out), for a final 52.0%, ranked 0 of the week for confidence. Market side: the other side.
What it teaches: We disagreed with the market and lost. This is our error, not an upset. The injury adjustment pushed us 52.0 points off the market's number and it was the wrong direction -- exactly the case for weighting injuries less, not more.
Improvements Under Consideration (10)
Every idea carries what would settle it. Ideas that were tested and failed stay on this list rather than being deleted — the record of what doesn't work is what stops the same idea being re-proposed every month.
Being tested
Grades plus injuries against the spread
Unit grades and injury detail together reached 53.11% ATS over 1,384 walk-forward games (z=2.31), which is above the 52.4% break-even. It is the only one of six feature sets to cross the bar.
How we'd settle it: Treat as UNPROVEN and re-test on a full fresh season. Two reasons for caution: six sets were tested, so one crossing z=2 is close to what multiple comparisons produce on their own; and the superset "everything" scored 51.52%, which is not how a real signal behaves -- adding information should not destroy it. The standing rule applies: a candidate must repeat on data it has never seen before it counts.
Measured: 53.11% ATS, z=2.31, +0.71pp over break-even, 1,384 held-out games (+0.71pp)
Being tested
Revise picks against late-breaking information
The edge is in the window, not the weights. Every backtest trained on CLOSING lines, which already contain late injury, weather and line movement -- which is exactly why injuries measured as worthless. The line we hold on Tuesday does not contain them. Re-checking an hour before the deadline and revising uses information the Tuesday line did not have, and is the only place the evidence says an edge could still exist.
How we'd settle it: Record the inputs each pick was made against, re-check them before the deadline, and log every revision with what moved. After enough weeks, compare the accuracy of original picks against revised ones on the games where the re-check actually changed something. That is a direct measurement, not an argument.
Proposed
Snapshot line movement instead of only the closing line
Every measurement so far trains on nflverse historical CLOSING lines, which already contain late injury and weather news. The line we actually hold for an upcoming game is posted days earlier. So "the market prices injuries" is proven for closing lines and UNPROVEN for the early line we use -- an injury breaking after our line was posted is the single most plausible place a real edge exists.
How we'd settle it: Record the spread, total and moneyline on every build from now on, building a real history of how each line moves. Then measure whether news arriving between the early line and close is predictable from our injury feed. Needs weeks of collection before it can answer anything.
Proposed
Give the model room to disagree with the market
The first edge search found the model takes a different side on only 1.6% of games. An edge is impossible without disagreement -- the current model is close to a re-statement of the line. Either it needs features carrying real independent weight, or it should be described as a market tracker rather than a model.
How we'd settle it: Measure how the disagreement rate and the win rate on disagreements move together as features are added. More disagreement is only progress if the disagreements are right.
Proposed
Use the player and unit grades as model inputs
Grades exist for 16,243 player-seasons and 791 units, including a real offensive-line grade, and none of them feed the pick model. A line-quality mismatch or a graded starter being out is plausibly information the market prices imperfectly.
How we'd settle it: Add unit-grade differentials and a graded-starters-out term to the walk-forward backtest. Adopt only if it beats the market on disagreements at n>=30 and z>2, per the standing bar.
Proposed
Weather at kickoff, not just the venue
Stadium coordinates and a free National Weather Service feed are already wired up for the site, but no weather value reaches the model. Wind in particular has a well-documented effect on scoring that a line posted days earlier cannot fully price.
How we'd settle it: Attach forecast wind, temperature and precipitation at kickoff to each game, backfilled where history allows, and test as a feature -- with particular attention to totals rather than winners.
Proposed
Report the disagreement rate on every pick set
A model that agrees with the market on every game should say so plainly, rather than presenting itself as an independent opinion. This is a disclosure fix, not an accuracy fix.
How we'd settle it: No test needed. Show, on the picks page and in the email, how many of the week's picks differ from the market side.
Proposed
Per-game deadlines, not one weekly deadline
Yahoo's deadline is five minutes before EACH game, so a single weekly cutoff is wrong. The current "final" run keys off the week's first kickoff, which is right for the earliest game and increasingly stale for a Sunday afternoon or Monday night one.
How we'd settle it: Schedule the final re-check per game against that game's own kickoff, keeping the owner's chosen one-hour safety margin ahead of Yahoo's five-minute cutoff.
Adopted
Pick against the spread, not straight up
The pool scores ATS with no confidence points. Straight-up accuracy of 67% is irrelevant to it: what matters is beating the spread, where 50% is the null and roughly 52.4% is break-even against standard juice. Every pick shipped so far answered the wrong question for this pool.
How we'd settle it: Already measured. Walk-forward 2021-2025, 1,384 held-out games: spread only 51.37%, unit grades 51.88%, grade matchups 51.16%, grades+injuries 53.11%, everything 51.52%.
Measured: ATS is the target; straight-up picks do not serve this pool
Blocked
Read the spread from Yahoo, never from a betting feed
Yahoo posts its own line for the pool, and it need not match nflverse's closing line. Picking ATS against a different number than the pool grades against is simply wrong, however good the model is -- a half-point difference flips games.
How we'd settle it: Scrape the pool page's own spread before every prediction run and store it per game. The model must REFUSE to emit an ATS pick when Yahoo's line is unavailable, rather than substituting a betting-market line that looks close.
Measured: blocked: needs a fresh Yahoo login, which requires the owner's 2FA code in real time
How This Stays Honest
Predictions are stored before kickoff and never rewritten. Grading is deterministic and generated from the actual feature values — no language model writes these explanations, because a reason you cannot check is worse than no reason. Every performance figure is walk-forward. The edge bar (30+ disagreements, z beyond 2) was set before any result was seen, so it cannot be moved to make a finding look better. Run by scripts/model/review-picks.mjs and scripts/model/find-edge.mjs on every build.