NFL Week 3 AI Predictions: Hiding the Betting Market Beat Kalshi
Every week this NFL season, we're pitting 8 AI agents — 4 with Daggy's statistical engine and 4 without — against Kalshi's prediction market. Here's the Week 3 recap.
NFL Week 3 AI Predictions vs Kalshi: Brier Scores
Our agents beat Kalshi's market predictions in Week 3 with a Brier score of 0.243 vs 0.258. The Broncos' tight win over the Rams on Sunday night led Daggy to win the head-to-head at 9-7 vs Kalshi's 8-8.
This week, we redacted market information from our agents because they had previously anchored too heavily on the market price. This led our predictions to move much farther from Kalshi's. Our forecast averaged 3.7 points from Kalshi's predictions compared with 1.3 points in Week 1 and 1.9 points in Week 2 (before extremization). The Rams versus Broncos was the first game all season where our system picked a different winner than Kalshi.
- This was Kalshi's worst week of the season, picking just 8 of 16 winners. Going back to 2006, betting favorites won half or fewer of their games in about 12% of the weeks.
- Home teams were exceptionally strong this week, winning 11 of 16 games (69%) vs a ~56% base rate. Home Elo underdogs won 7 of 9 games (78%), against a historical 43%.
- Market-redacted agents with Daggy's statistical engine picked 61% correctly compared with 55% for agents without.
How the Agents Forecast
We used gpt-6-luna to remove market information from our agents' search results. This led the agents to disagree far more. Agents without this market redactor — which we ran this week as a baseline — averaged 1.4 points from Kalshi's predictions versus 4.9 points for our market-redacted agents. The per-game spread (highest minus lowest) for the market-aware agents was 3.7 points versus 11.9 points for the market-redacted agents.
Agents with market information typically overrode any statistical models they built. Here is Kimi, with market information and Daggy's engine, on the Rams vs Broncos:
MARKET (strongest signal): … Betting lines are the best-calibrated NFL predictor and already incorporate form, injuries, rest, and altitude … ELO MODEL (pulls the other way): Pre-game Elos are nearly even (DEN 1588.9 vs LA 1591.1). A linear home_win model on Elo + rest (trained on 2000–2026w2) predicts DEN 0.567 [90% PI 0.538–0.600], driven by home field and Denver's 7-vs-6 rest-day edge … Synthesis: blend of market (~0.444) and Elo model (~0.567) weighted toward the market: 0.46.
This shows that although its own model disagreed with the market by 12 points, Kimi moved only 1.5 points away from the market. It justified this anchoring by calling betting lines the "best-calibrated NFL predictor." Contrast this with the market-redacted version of Kimi on the same game:
With even Elo, home-field advantage … plus a small rest edge … drives a trained logistic model (elo_diff, rest_diff, same_division, fit on 2015-2025) to 0.563 for DEN … Injuries break the tie toward Denver … Rams star WR Puka Nacua missed both Wednesday and Thursday practices … Net: ~0.54 base + even-Elo HFA/rest (+~2%) + injury edge (+~2%) ≈ 0.58.
Redacting market information led this Kimi agent to use what it found in search, such as injuries, to become more bullish on Denver than its statistical models suggested. With market information, the agents largely ignored the statistical models because, as Kimi noted, "Betting lines are the best-calibrated NFL predictor." In 7 of 16 games, for instance, at least one market-aware agent moved at least 9 points from its own model toward Kalshi.
Because the agents weighed the market so heavily, Daggy's statistical engine had no measurable impact on their predictions. But removing the market signal led agents to rely much more heavily on the engine. The tables below show that agents with access to it picked one more winner and had a lower Brier score than their twins.
| Opus | Gemini | GPT | Kimi | Total | |
|---|---|---|---|---|---|
| Redacted, with Daggy engine | 9-7 | 10-6 | 10-6 | 10-6 | 39-25 (61%) |
| Redacted, without Daggy engine | 8-8 | 9-7 | 9-7 | 9-7 | 35-29 (55%) |
| Market (with and without engine) | 8-8 | 8-8 | 8-8 | 8-8 | 32-32 (50%) |
| Opus | Gemini | GPT | Kimi | Avg of agents | Averaged forecast | |
|---|---|---|---|---|---|---|
| Redacted, with Daggy engine | 0.254 | 0.237 | 0.242 | 0.228 | 0.240 | 0.240 |
| Redacted, without Daggy | 0.260 | 0.248 | 0.247 | 0.246 | 0.250 | 0.249 |
| Market, with Daggy | 0.252 | 0.257 | 0.251 | 0.251 | 0.253 | 0.253 |
| Market, no Daggy | 0.255 | 0.256 | 0.253 | 0.249 | 0.253 | 0.253 |
Changes for Week 4
For Week 4, we will continue using the market-redacted agents and retire the market-aware ones. To observe the market's impact, however, we will run two versions of the supervisor agent, one with access to the market and one without.
We are also incorporating memory and reflection for our Week 4 agents. This will allow them to learn from and improve upon their previous forecasts. Memory may also push the agents further apart throughout the season.