中文版

Which AI predicted the 2026 FIFA World Cup best?

WC2026-Agents: four LLM forecasting agents versus the betting market on all 104 matches of the 2026 FIFA World Cup (11 June to 19 July 2026). By Jiacheng Ding and Cong Guo, University of Memphis.

Short answer. Claude Opus 4.8, ChatGPT (GPT-5.5), Gemini 3.1 Pro and Grok (Expert Mode) each forecast and placed a virtual bet on every 2026 World Cup match about a day before kickoff. They picked the same outcome in 96 of 104 matches and reached 65.4% to 68.3% accuracy. None beat the betting market's Brier score (0.469). Where they differed was money: at real pre-match odds Grok made the most virtual profit (+$650, ROI +10.3%), Gemini +$322, ChatGPT +$118, and Claude lost $275. Simply backing the market favourite every match made +$1,041.

Results at a glance

ForecasterAccuracyBrier (lower is better)Log-lossVirtual profitROIBets
Betting market (vig-free odds)68.3%0.4690.807+$1,041*104
Grok (Expert Mode)68.3%0.47060.803+$650+10.3%104
Gemini 3.1 Pro65.4%0.48280.820+$322+3.7%103
ChatGPT (GPT-5.5, high reasoning)68.3%0.47290.814+$118+8.0%55
Claude Opus 4.866.3%0.47050.806-$275-18.1%73

Accuracy is the share of 104 matches where the top-probability outcome (home win, draw, away win after 90 minutes) happened. *Flat stake on the market favourite in every match. Agents chose their own stake (up to $100) or could skip a bet.

What is WC2026-Agents?

WC2026-Agents is an open benchmark and dataset for evaluating large language models as autonomous forecasting agents on real future events. For each of the 104 matches of the 2026 FIFA World Cup, four frontier models ran the same search, act and reflect loop: search the web, commit to win/draw/win probabilities and a virtual bet of up to $100, and after the match reflect on the result given only the final score. Every match kicked off after the models' training cutoffs, so the benchmark is contamination-free by construction. The pre-match betting market is scored as a fifth forecaster.

The release contains 416 forecasts with verbatim reasoning, 414 post-match reflections, ground-truth results including penalty shootouts, pre-match 1X2 odds with sources, and a reproducible evaluation pipeline.

Key findings

1. The four AIs make almost identical predictions

They chose the same outcome in 96 of 104 matches (92%). The 8 disagreements were all 3 against 1, with Gemini the lone dissenter in 6. Accuracy spans only three matches: ChatGPT and Grok 68.3%, Claude 66.3%, Gemini 65.4%. In the 32 knockout matches all four scored exactly 71.9%.

Heatmap of pairwise same-pick rates between Claude, ChatGPT, Gemini, Grok and the betting market, and a bar chart showing most matches were unanimous
How often each pair of forecasters picked the same outcome, and how many matches were unanimous.

2. No AI beat the betting market

The vig-free market probabilities scored a Brier of 0.469, better than every agent (0.4705 to 0.4828). On log-loss Grok (0.803) and Claude (0.806) sat marginally below the market (0.807), too close to call over 104 matches. A flat bet on the market favourite every match returned +$1,041, more than any agent earned.

3. Betting separates the models

Same predictions, different money. Grok bet on all 104 matches and won 70 bets for +$650 (ROI +10.3%), including +$595 and 24 of 32 winning bets in the knockout stage. Gemini staked the most ($8,660) for +$322. ChatGPT was the most selective (55 bets) for +$118 (ROI +8.0%). Claude won only 24 of 73 bets and lost $275 (ROI -18.1%).

Line chart of cumulative virtual betting profit over 104 matches: market favourite +1041 dollars, Grok +650, Gemini +322, ChatGPT +118, Claude -275
Cumulative virtual profit at real pre-match odds, group stage then knockout.

4. Betting against the market mostly lost

Claude backed an outcome other than the market favourite on 58% of its bets, ChatGPT 36%, Gemini 14%, Grok 5%. Those contrarian bets won 21% to 40% of the time, versus 48% to 69% when agreeing with the market, and lost money for Claude (-$167), Gemini (-$112) and ChatGPT (-$38). Grok's 5 contrarian bets made +$64.

Bar charts of each AI's share of bets against the market favourite and the profit from market-conforming versus contrarian bets
Share of bets against the market favourite, and profit split by conforming versus contrarian bets.

5. They read different evidence

Share of forecasts whose reasoning cites betting odds or the market: Claude 100%, ChatGPT 92%, Grok 63%, Gemini 12%. Grok most often cited squad quality (88%) and recent form (76%). Head-to-head records (4% to 23%) and fatigue or rest (7% to 10%) were rarely mentioned by any model.

Dumbbell chart of how often each AI cites 13 reasoning factors such as market odds, squad quality, form, injuries and FIFA ranking
How often each agent cites each factor in its pre-match reasoning.

6. Draws are the blind spot

24 of 104 matches (23%) were level after 90 minutes, but no model made a draw its top pick more than 4 times (Claude 1, ChatGPT 2, Grok 2, Gemini 4). Of the 32 matches all four agents got wrong, 22 were draws.

7. Upsets that fooled every AI

In the knockout stage all four agents missed 8 matches: Germany 1-1 Paraguay, Netherlands 1-1 Morocco, Australia 1-1 Egypt and Switzerland 0-0 Colombia (decided on penalties), plus Brazil 1-2 Norway, USA 1-4 Belgium, France 0-2 Spain and France 4-6 England. In their post-mortems the agents most often named the underdog's defensive organisation as the factor they missed (62% of wrong picks) and the favourite's squad quality as the factor they over-weighted (40%).

Dot plot of the eight 2026 World Cup knockout matches that all four AIs mispredicted, with the consensus probability on the favourite
Knockout matches every agent got wrong, with the average probability they gave the favourite.

8. AIs differ in admitting mistakes

After a wrong pick, each agent labelled its own forecast. Gemini called it incorrect 86% of the time, Grok 61%, Claude 49% and ChatGPT 36%; the rest were described as partially correct or correct.

Stacked bars of how each AI labels its own wrong predictions: owns the error, partially correct, or denies
Self-labels on wrong picks in post-match reflection.

How the experiment worked

FAQ

Which AI predicted the 2026 FIFA World Cup best?

It depends on the metric. In the WC2026-Agents benchmark (all 104 matches), ChatGPT (GPT-5.5) and Grok (Expert Mode) tied for the highest accuracy at 68.3%, Claude Opus 4.8 (Brier 0.4705) and Grok (0.4706) had the best probability scores, and Grok earned the most in virtual betting (+$650, ROI +10.3%). None of the four beat the pre-match betting market's Brier score of 0.469.

Did any AI beat the betting odds at the 2026 World Cup?

No, not convincingly. The vig-free market probabilities scored a Brier of 0.469, better than all four agents (0.4705 to 0.4828). On log-loss Grok (0.803) and Claude (0.806) were marginally below the market (0.807), a gap too small to be meaningful over 104 matches. A flat bet on the market favourite in every match earned +$1,041, more than any agent.

Which AI made the most money betting on the 2026 World Cup?

Grok (Expert Mode): +$650 virtual profit on 104 bets (ROI +10.3%), including +$595 in the knockout stage. Gemini 3.1 Pro made +$322, ChatGPT (GPT-5.5) +$118, and Claude Opus 4.8 lost $275 (ROI -18.1%). Bets were virtual, sized by each model (up to $100) and settled at real pre-match odds.

How accurate were ChatGPT, Claude, Gemini and Grok at predicting World Cup 2026 matches?

Over all 104 matches (win/draw/win after 90 minutes): ChatGPT 68.3%, Grok 68.3%, Claude 66.3%, Gemini 65.4%. The betting market's favourite was also right 68.3% of the time. In the 32 knockout matches all four agents scored exactly 71.9%. The four models chose the same outcome in 96 of 104 matches.

Did AI predict Spain winning the 2026 World Cup final?

Before Spain beat Argentina 1-0 in the final on 19 July 2026, all four agents rated a Spain win at least as likely as any other outcome, with Spain win probabilities of 35% to 44%. The agents forecast each match separately about a day before kickoff; they did not make one pre-tournament champion pick.

Which 2026 World Cup matches did every AI get wrong?

All four agents missed 32 of the 104 matches, and 22 of those were level after 90 minutes. In the knockout stage all four missed 8: Germany 1-1 Paraguay, Netherlands 1-1 Morocco, Australia 1-1 Egypt and Switzerland 0-0 Colombia (all decided on penalties), plus Brazil 1-2 Norway, USA 1-4 Belgium, France 0-2 Spain and France 4-6 England.

Why do AI models fail to predict draws in football?

In WC2026-Agents, 24 of 104 matches (23%) were draws after 90 minutes, yet each model made a draw its single most likely outcome in at most 4 matches. The models gave draws a sensible average probability (22.7% to 23.8%), but a draw is rarely the single most likely result, so top-pick accuracy misses almost every draw.

Does it pay for an AI to bet against the betting market?

Mostly not. Bets on an outcome other than the market favourite won 21% to 40% of the time, versus 48% to 69% for bets that agreed with the market. Contrarian bets lost money for Claude (-$167 on 42 bets), Gemini (-$112) and ChatGPT (-$38). Grok rarely went against the market; its 5 contrarian bets made +$64.

Why is WC2026-Agents contamination-free?

Every match kicked off after the models' training cutoffs, so the results could not be in any model's training data. The same four model versions were used unchanged from 11 June to 19 July 2026, and every forecast was recorded before kickoff.

Where can I download the WC2026-Agents dataset?

It is free under CC BY 4.0 on Hugging Face (dingjiacheng/wc2026-agents) and GitHub (graphuofm/FIFA2026LLM), with code under MIT. It contains 416 forecasts, 414 post-match reflections, results including penalty shootouts, and pre-match odds for all 104 matches.

Who created WC2026-Agents?

Jiacheng Ding and Cong Guo of the University of Memphis. The benchmark is described in the arXiv paper "FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches" (arXiv:2607.17765).

Use the data

from datasets import load_dataset
forecasts = load_dataset("dingjiacheng/wc2026-agents", "forecasts")
odds = load_dataset("dingjiacheng/wc2026-agents", "odds")

Reproduce every table and figure: pip install -r requirements.txt then python src/run_all.py in the GitHub repository.

Cite

@misc{ding2026wc2026agents,
  title         = {{FIFA} World Cup 2026 as a Contamination-Free Benchmark for
                   {LLM} Forecasting Agents: Four Models, a Bookmaker, and 104 Matches},
  author        = {Ding, Jiacheng and Guo, Cong},
  year          = {2026},
  eprint        = {2607.17765},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2607.17765}
}