The Roundup
October 6, 2026

NFL Week 4 AI Predictions: Daggy Wins Two in a Row

Every week this NFL season, we're pitting 8 AI agents — 4 with Daggy's statistical engine and 4 without — against Kalshi's prediction market. Here's the Week 4 recap.

NFL Week 4 AI Predictions vs Kalshi: Brier Scores

For the second straight week, our agents beat Kalshi's market predictions in Week 4 with a Brier score of 0.214 vs 0.223 (a lower Brier score is better). Daggy went 11-5 picking winners vs Kalshi's 9-7. For the season, Daggy's 0.230 Brier score barely ekes out Kalshi's 0.231, or effectively tied. Daggy is 43-21 (67%) picking winners vs Kalshi's 40-24 (62.5%).

Week 4 · 16 games
Daggy
11-5
68.8%
Brier
0.214
Kalshi
9-7
56.3%
Brier
0.223
Lower Brier is better. 0 is a crystal ball, 0.25 a coin flip.
  • Favorites continued to struggle: Elo favorites predicted 8 of 16 games (50%) vs a historical rate of 63.7% and Kalshi favorites predicted 9 of 16 games (56%) vs a typical market win rate of 66.7%.
  • This was Kalshi's best week by Brier score since Week 1 and Daggy's lowest of the season.
  • The market-aware supervisor again overly anchored on market odds.

How the Agents Forecast

Week 4 was the first week agents used memory and reflections. After Week 3, agents ran a post-mortem to describe what went well and what went wrong. They kept a card for each team and a card of forecasting lessons that they can update throughout the season. For example, in Week 3, Kimi with Daggy's engine noted the following lesson in its post-mortem:

"When Elo, the trained model, form, rest, and injuries all converge, trust the number and don't dilute it with narrative hedging."

Then in this week's LAC@SEA prediction it used this lesson to move from 0.76 on SEA to 0.80, stating: "per PM [post-mortem] 2 not diluting toward 0.50 when everything converges." Seattle won and the move improved its Brier score. But some rules pulled the agents toward 50-50, as Opus wrote "near pick'em (Elo+HFA ~45-58%), no QB change: stay within 45-55%." More than a third of the agents' predictions this week fell between 45% and 55%, compared to zero games for Kalshi. Agents may have reason to be more aggressive, such as other game-specific or personnel factors, and in September Elo mostly represents the previous season.

Agents with Daggy's statistical engine typically outperformed those without. Gemini with Daggy's engine had a 0.207 Brier score while its engine-free twin had 0.225. Only GPT without Daggy did better than its twin with a 0.208 Brier score compared to 0.210.

OpusGeminiGPTKimiTotal
With Daggy engine10-611-59-6-111-541-22-1 (64%)
Without Daggy engine9-79-711-511-4-140-23-1 (63%)
BrierOpusGeminiGPTKimiAvg of agentsAveraged forecast
With engine0.2180.2070.2100.2080.2110.210
Without engine0.2400.2250.2080.2130.2210.220

Finally, the market-aware supervisor agent, which we didn't use in our forecast, again anchored heavily on the market odds it found in search and downweighted the agents' analysis. The market-redacted supervisor averaged 5.5 points from Kalshi while the market aware supervisor averaged 1.7 points from Kalshi. For example, for the ARI@NYG game the agents typically favored NYG while the market favored ARI, but the market-aware supervisor wrote:

No run had the betting line, and it settles which side is favoured … The market already prices in these losses. Runs starting from Elo plus home field (0.58–0.62) overrated NYG …

It then adjusted the agents' average of 52% NYG to 52% ARI, which is the midpoint to Kalshi's 56% ARI. The agents with 0.58-0.62 predictions for New York were all those with the engine-on agents, which the supervisor noted were overrating NYG. New York won.

Changes for Week 5

In Week 5 we will continue with memory and allow the agents to adjust their views of the teams and their best practices for prediction. We added one sentence to the prompt: "Forecasts are scored by the Brier score." We did this because we want the agents to report their true beliefs, and not to hedge by following rules that force the agent to stay within "45% and 55%" for example. Because the Brier score is a strictly proper scoring rule, the agents' best strategy is to report their true probabilities