I built an expected-points model on 1.8M plays (2015–2025, next-scoring-event construction, the half as the horizon) and graded every fourth-down decision in 1,430 games against it. 22,218 decisions.
Coaches match the model 81.3% of the time, which is better than the genre usually implies. The 18.7% they miss is the interesting part, and it has drifted steadily in one direction.
The drift is real
| Season |
Over-aggressive |
Too passive |
Share aggressive |
| 2015 |
213 |
181 |
54.1% |
| 2018 |
198 |
141 |
58.4% |
| 2021 |
246 |
132 |
65.1% |
| 2024 |
246 |
113 |
68.5% |
| 2025 |
232 |
130 |
64.1% |
That's +1.05 percentage points a year, t = 3.53. In 2015 a coach's fourth-down errors were close to a coin flip between too bold and too timid. Today two-thirds of them are boldness. The sport was told to stop punting, and it did.
The obvious read is that coaches have overcorrected. I think that read is wrong, and the reason is the most interesting thing in the data.
The over-aggression is entirely underdogs
Split every decision by what the betting market thought of the team making it:
| Market position |
Went for it |
Model says |
Gap |
| Favorite by 14–21 |
21.9% |
22.7% |
−0.8 |
| Favorite by 7–14 |
23.6% |
23.3% |
+0.4 |
| Underdog by 3–7 |
25.6% |
19.9% |
+5.7 |
| Underdog by 7–14 |
24.1% |
16.6% |
+7.5 |
| Underdog by 21+ |
21.6% |
15.1% |
+6.5 |
Favorites are already playing the model almost exactly — 53.6% of their mistakes are over-aggression, essentially balanced. Among underdogs it's 67.4%.
So: are the underdogs wrong? You can test it directly. Take every situation where the model said don't go, and compare the teams that went anyway against teams in the same spread band and the same score state that obeyed.
- Favorites who went anyway: 54.9% win rate vs 59.8% for obeying — −4.8 points, 95% CI [−7.0, −2.5]
- Underdogs who went anyway: 13.1% vs 13.1% — −0.0 points, 95% CI [−1.4, +1.1]
For a favorite, the expected-points model is right about winning too, and disobeying it costs about five points of win probability. For an underdog it costs nothing, and that null is tight rather than empty: 1,319 decisions, with an interval that rules out any effect bigger than about a point and a half either way.
That's not a failure of the coaches. It's a failure of the yardstick. Expected points is risk-neutral. A team that is going to lose seven times in eight does not want the highest average score — it wants the widest distribution, because it needs the tail. Trading expected points for variance is the correct play there, and a risk-neutral model records it as an error every single time.
Which inverts the lecture. The coaches being scolded for aggression are the ones for whom aggression is free. The coaches for whom the model is genuinely right — favorites, who have edge to protect and want to convert it into certainty — are the ones already following it, and the ones who pay when they don't.
The expensive decisions are the ones that look like they worked
This part survives regardless of which yardstick you prefer. Sort every wrong call by what the viewer saw:
| What happened |
Decisions |
EP lost |
Share |
| Gambled and failed — visible |
1,265 |
−1,064 |
39.6% |
| Gambled and converted — invisible |
1,197 |
−880 |
32.8% |
| Kicked and made it — invisible |
717 |
−367 |
13.7% |
| Punted — invisible |
751 |
−281 |
10.5% |
| Kicked and missed — visible |
226 |
−93 |
3.5% |
Fifty-seven percent of the damage comes from decisions where the play worked. The fourth-down gamble that fails on national television is 40% of the cost and 100% of the discourse.
The purest case is the short field goal. There were 912 decisions where a team kicked and the model said go — average situation fourth-and-4 from the opponent's 17, and 399 of them from inside the opponent's 10. 77% of those kicks were good. Three points went on the board, the kicker got a pat, the broadcast moved on, and the decision cost the team expected points anyway.
A made field goal is the most expensive thing in football that nobody complains about.
Three games that turned on it, none of them the way you'd remember
Utah at Colorado, 2016. Utah lost by five. They kicked three chip shots:
Q3 2:58 tied 4th & 3 at the 3 kicked -1.12 20-yard FG good
Q3 11:54 down 3 4th & 4 at the 4 kicked -0.77 21-yard FG good
Q4 13:53 down 4 4th & 5 at the 5 kicked -0.45 22-yard FG good
Three made kicks from inside the five, worth 2.34 expected points against going for it, in a five-point loss. Every one of them appeared in the box score as three points and a success.
Kansas at TCU, 2016. Kansas lost by one. The remembered play is a desperate fourth-and-22 from their own 31 with forty seconds left — which the model scores at −2.85 and which is, in win-probability terms, obviously correct, because punting there is conceding. The decision that actually cost them came in the third quarter, leading by two: fourth-and-4 at the TCU 4, and they kicked. Made it. −0.77. In a one-point game.
Penn State at Ohio State, 2017. Lost by one. Same shape. The fourth-and-15 gamble in the last ninety seconds is the play everyone argued about, and it's the one the model penalizes hardest (−3.16) and is least qualified to judge. The quiet one — fourth-and-6 at the Ohio State 6, up eleven, kick — cost 0.18, and nobody has mentioned it since.
Across the corpus, 25 of 673 one- and two-score losses (3.7%) featured a team whose own fourth-down decisions cost more expected points than it lost by.
Method
Expected points by game state comes from the next-scoring-event construction over 1.8M plays, 2015–2025, with the half as the horizon. Near-goal EP is +5.10 and own-10 EP is −0.81, field-position correlation −0.981.
Conversion probability is fit on third downs rather than fourth, because teams choose to go for it when they expect to convert and fourth-down rates are selected upward. Field-goal probability is a monotone logistic fit by kick distance. Punt value is taken from where the opponent's next drive actually started. Win-rate comparisons are within spread band and score state; intervals are bootstrapped clustered on game.
What it doesn't do. It doesn't grade play-calling — run vs pass is an equilibrium problem, and a model that ignores the defense's response will always conclude everyone should pass more. It doesn't weight expected points by leverage, which is why the costliest decision of a given week is often one taken in a blowout. And, as the central finding says plainly, it doesn't know what a team's objective function is. The underdog result is a demonstration of what happens when you assume it.
Objections worth answering up front
- "Your model just doesn't understand game situations." Correct, and that's the finding. It's risk-neutral, and the underdog result is the demonstration of what that costs you as an analyst.
- "Fourth-down conversion rates are biased because coaches pick their spots." Yes. That's why conversion probability is fit on third downs.
- "Small sample." 22,218 decisions, 1,430 games, eleven seasons, intervals bootstrapped clustered on game.
- "You're just describing garbage time and desperation." Score state is controlled, and comparisons are within spread band and score state.
Happy to share the per-season and per-band numbers if anyone wants to poke at them.