AI Predictions vs The Market Week 1
Every week this season, we're pitting 8 AI agents — 4 with Daggy's statistical engine and 4 without — against Kalshi's prediction market. Here's how our Week 1 NFL predictions went.
The Week 1 Scoreboard
Kalshi beat our agents in Week 1 of the NFL regular season with a Brier score of 0.211 compared to our 0.216 (a lower Brier score is better). Both Kalshi and our system predicted 12 out of the 16 games correct. Among the agents, Kimi K2.5 with Daggy had the lowest Brier score of the week (0.210) and GPT-5.5 with Daggy had the highest (0.224).
- The agents relied heavily on market data. After finding market odds, agents moved their forecasts by an average of 6.3 percentage points.
- Agents with Daggy access ran statistical models on games with closer to even odds, suggesting they reached for tools on games that were harder to predict.
- Extremizing the agents' forecasts, or pulling them away from 50-50 using standard methods, would have cut our Brier from 0.216 to 0.209.
How the Agents Reasoned
Most agents began their runs by searching for the market's odds. GPT-5.5 without Daggy and opus-5 in both arms opened every run by searching for market odds. Kimi-k2.5, despite using search the most, asked for market odds about 20 percent of the time and only after it had already run several other searches about the game. The models pulled odds from oddsmakers like BetMGM and prediction markets like Kalshi. Papers, such as AIA Forecaster and Bayesian Linguistic Forecaster have found that market information can allow LLM forecasters to substantially improve their predictions, and these agents all follow suit.
The models leaned on base rates. Nearly all noted that home teams win about 55 percent of games. Claude opus-5 found a more granular base rate by using Daggy to find that the home team only won "~50.3% of Week 1 games (nfl_games query). So I started near 0.50-0.52 for a generic home team." The agents, therefore, followed the advice from Tetlock and Gardner's Superforecasting by incorporating base rates into all of their forecasts. Opus-5's analysis also shows that base rates have many cuts as it considered both the overall home win percentage and the Week 1 home win rate.
The agents with Daggy access built statistical models for games that were more difficult to call. The market odds sat closer to a tossup for games when the agents built models and farther from it on games they skipped modeling. The market also struggled to predict these games, as Kalshi's scored 0.231 on the modeled games and 0.165 on the others.
Typical models relied on features like home win percentage and previous-season point differentials. But in Week 1 those models used last year's rosters. The models saw this. As, gpt-5.5 noted for Giants at Cowboys "The prior-season model is relatively simple and does not know current QB depth, coaching, roster changes, or matchup-specific tactics". But they didn't fully discount historical performance either, sometimes investigating it directly. Opus-0 on Cleveland at Jacksonville noted "Among 2001-2025 REG Week 1 games where the home team's prior-season win pct was ≥0.70 and the away team's was ≤0.40, the home team won 14 of 20 (70%). That small sample suggests ~70%, but it lumps in many home teams favored by far less than 8 points and ignores current-season information." It then upped its belief from 0.68 to 0.75. The agents recognized the data was thin this early in the season. We expect them to use more sophisticated models once the data better describes the 2026 versions of each team.
Changes for Week 2
For week 2, we are adding Platt scaling to extremize the agents' forecasts, as this would have helped all eight agents and is recommended in the literature. Furthermore, we are giving the Daggy agents Elo rankings built with 538's published parameters.