NFL AI Predictions vs The Market: Week 2
Every week this NFL season, we're pitting 8 AI agents — 4 with Daggy's statistical engine and 4 without — against Kalshi's prediction market. Here's the recap on our Week 2 NFL predictions.
AI Agents vs Kalshi: Week 2 Scoreboard
We made our agents more aggressive with extremization, and paid the price. Kalshi beat our agents in Week 2 of the regular season with a Brier score of 0.231 compared to Daggy's 0.248 (a lower Brier score is better). If we had avoided extremization and took a simple mean, Daggy's Brier score would have tied Kalshi with 0.231 — shoulda, coulda, woulda. Both Kalshi and our system correctly predicted 11 out of the 16 games.
- Although both Daggy and Kalshi struggled this week, these weren't historic misses. Going back to 2006, the market scored 0.227 or worse in about 32% of the weeks and 0.248 or worse in about 19%.
- Agents with access to Daggy's data-analysis tools paid close attention to historical baselines and Elo rankings, neither of which predicted well this week. Home teams won 7 of 16 games (44%) against a 55% baseline. Home Elo favorites won 5 of 10 games (50%) against a historical 69%.
- Extremization generated a few big misses and hurt Daggy overall, likely because the agents anchored on market odds, leaving their predictions highly correlated.
How the Agents Forecast
Like Week 1, all agents searched for market information and base rates, then refined these to make their predictions. Our best game of the week shows the process. Kimi with Daggy began its Packers at Jets forecast by finding Week 2 home win rates, which were 56.8% and favored the Jets. After noting Green Bay had a +112 Elo advantage, it adjusted toward favoring the Packers. It then built linear models and neural networks and weighed the model output against market odds before finally settling on a 64% probability for the Packers versus Kalshi's 61%. After extremizing the forecasts, we predicted 73% on GB, which paid off despite the game going to overtime.
Daggy's biggest misses came from Cleveland at Tampa Bay and New Orleans at Baltimore. Both home teams were market favorites with higher Elo ratings. Extremization pushed Daggy's prediction for Tampa Bay from 77% to 89% and Baltimore from 75% to 87%; Kalshi had both at 79%. Although Daggy and Kalshi missed both games, Daggy was more confident. This confidence cost Daggy 0.0189 more in Brier than Kalshi, which was more than the 0.017 total difference between them for the week.
This raises questions about extremization. When forecasters share information, or their forecasts are correlated, we should extremize less than when their forecasts are independent. We relied on values from Neyman and Roughgarden and Alur et al. to choose the extremization parameter, which determines how aggressively to pull our forecasts from 0.5. But this still may have been too aggressive. Although Alur et al. analyzed groups of LLMs making predictions based on correlated information, our agents may have had higher correlation because of the strong information in market signals.
Changes for Week 3
Extremization was probably too aggressive here because of the agents' information overlap. Following research from Bridgewater, we will start acting on our supervisor agent, which we've had running in the background since Week 1. The supervisor reconciles the other agents' forecasts and outputs a probability and a confidence score. If it is not confident, we will fall back to the agents' average.
Previous research showed that agents did better if they had a small LLM — Haiku in our case — summarize the search results. But given that the agents quickly anchored on the odds, causing their predictions to be highly correlated, we will give the agents the raw results from their search queries along with the summary. Although they will still see the odds, the snippets may add some diversity to their predictions.
We are also updating the models. OpenAI and Anthropic released gpt-6-sol and opus-5-5, respectively, which are both faster and cheaper than the current versions. We are also moving Moonshot AI from Kimi-2.5 to Kimi-3 and Google's Gemini Pro to Gemini Flash.