futuresearch evals
Want to run on our benchmarks? Please contact us at evals@futuresearch.ai.
How our forecasting and research agents perform on our own benchmarks, on live public leaderboards, and on real markets.
Bench to the Future 3 (BTF-3)
BTF-3 is the third edition of our pastcasting benchmark: 2,386 resolved forecasting questions — 2,015 binary and 371 numeric — researched and forecast against a frozen web corpus. The dataset is available at huggingface.co/datasets/BTF-2/BTF-3. Paper to follow.
BTF-3 Leaderboard
Evaluated: June–August 2026
All scores are on the Brier scale; lower is better, and the best score in each column is bolded.
Agent | Pooled score(n=2,386) | Binary(Brier, n=2,015) | Numeric(RPS, n=371) |
|---|---|---|---|
1FutureSearch SOTA*§ | 0.116 [0.109–0.123] | 0.112 [0.103–0.120] | 0.123 [0.112–0.133] |
2Claude Opus 5 (xhigh)§ | 0.120 [0.113–0.128] | 0.116 [0.107–0.125] | 0.128 [0.118–0.140] |
3Claude Opus 4.8 (xhigh) | 0.132 [0.125–0.139] | 0.130 [0.121–0.139] | 0.135 [0.124–0.146] |
4GPT-6 Astra (low) | 0.135 [0.129–0.141] | 0.137 [0.130–0.144] | 0.130 [0.120–0.142] |
5GPT-5.6 Sol (medium) | 0.137 [0.131–0.144] | 0.139 [0.131–0.146] | 0.134 [0.122–0.146] |
6Muse Spark 1.3 (xhigh) | 0.141 [0.134–0.148] | 0.136 [0.128–0.144] | 0.150 [0.138–0.162] |
7Claude Sonnet 5 (xhigh) | 0.142 [0.135–0.149] | 0.139 [0.131–0.148] | 0.146 [0.135–0.158] |
8GLM-5.3-Flash | 0.149 [0.142–0.156] | 0.146 [0.138–0.153] | 0.155 [0.142–0.169] |
9GLM-5.3 | 0.149 [0.143–0.155] | 0.147 [0.140–0.155] | 0.152 [0.141–0.164] |
Binary questions are scored by the Brier score (mean squared error of the forecast probability), numeric questions by a normalized ranked probability score (RPS), which generalizes the Brier score to distributional forecasts. The pooled score averages across all questions, counting each numeric question three times as much as a binary one (numeric forecasts are more informative per question).
Brackets are 95% confidence intervals, computed by percentile bootstrap (5,000 resamples of the question set).
* FutureSearch SOTA synthesizes forecasts from multiple FutureSearch agent runs.
§ Training cutoffs vary widely across this board, from June 2025 to May 2026, and the later ones sit close to our question snapshots (late April to May 2026). Claude Opus 5 has the latest of them, May 2026, so it went in with the most recent knowledge of the world. We found no leakage of resolutions, but a fresher cutoff may still be worth something. The FutureSearch SOTA row carries the same mark because Claude Opus 5 is one of the agent runs it synthesizes from.
How we checked for leakage
Every question in this benchmark is asked as of an anchor date in late April or May 2026 and resolves over the months that follow. Claude Opus 5 has the latest training cutoff on the board, May 2026, so it could in principle have seen outcomes it is asked to forecast. Before publishing the Claude Opus 5 row we ran the following checks on it.
Where each model's knowledge ends. Claude Opus 5, May 2026. GPT-6 Astra, April 2026. GPT-5.6 Sol, February 2026. Claude Opus 4.8 and Claude Sonnet 5, January 2026. Muse Spark 1.3, likely November 2025. GLM-5.3, likely July 2025. GLM-5.3-Flash, likely June 2025. The three dates marked “likely” are ours, not the vendors’: Z.ai and Meta publish no cutoff for these models, so we measured each one by asking it about events with known dates and requiring it to answer UNKNOWN rather than guess, then taking the last month it still got right. All three drew the line cleanly, naming long-settled facts while declining outcomes from after their cutoff. We did not rely on what a model says about itself: two of the three named dates months away from what we measured, and the third never answered the question.
Recall probes. We asked Claude Opus 5 forced-choice questions about hard-to-guess May 2026 events (the Kentucky Derby winner, the Eurovision winner, the PGA Championship, the Champions League final, Colombia's first-round result). It scored 0 of 5 on the events a May-trained model should know, and also missed April's Masters winner. Its recallable knowledge fades around March 2026, well before the question anchors.
Timing gradients. If Claude Opus 5 had absorbed outcomes, its advantage should concentrate on questions asked or resolved closest to its training window. The opposite holds: its edge over the other models is flat to slightly increasing for later anchor dates, and questions resolving within days of a late-May cutoff show less of its advantage than questions resolving six weeks later.
Confidence audit. A model that knows outcomes produces unusual numbers of near-certain forecasts. Claude Opus 5 made fewer forecasts at or beyond 98 percent certainty than most other models on the board, and there is no question where it was near-certain and correct while every other model disagreed.
Trace audit. In Claude Opus 5's largest wins over the field, every load-bearing source it cited was published before the question's anchor date and was available to all models through the shared research snapshots. Most of those wins were correctly skeptical "no" calls on deadline-driven questions, a research-skill signature rather than a foreknowledge one. A sample of its search queries showed no terms whose relevance only emerged after the anchor dates.
What we cannot rule out: because Claude Opus 5 was trained closer to the anchor dates, it knows the spring 2026 news landscape better, which can make its research more targeted even with no knowledge of outcomes. That recency advantage runs the length of the board, whose cutoffs span June 2025 to May 2026, and it is one reason to read a cost or accuracy gap between two models with their cutoffs in view.
CHAMPS KNOW strategic emphasis
Mean Borda score per dimension (rank 1 = 10 … rank 10 = 1); higher means the dimension is more prominent in the agent's rationales. The top 3 agents are shown by default — click a name to add or remove it.
Binary calibration
Do the probabilities mean what they say?
Ten probability bins. On the dashed line, things called 70% happened 70% of the time; above it is underconfident, below it overconfident. Marker size tracks forecasts per bin. Top 3 shown; click a name to toggle.
Best Brier per dollar
Accuracy against the agent’s average LLM cost per question
Cost is the LLM spend of the forecasting agent itself (the model under test plus the smaller models it calls for sub-steps), averaged over the forecasts it made on the August 2026 cohort of the benchmark (about 600 questions), including refused and not-yet-resolved ones. That cohort is where every configuration has complete per-question billing, and every one ran it within the same ten days, so the points are directly comparable; accuracy on the other axis is still the full June to August score. Cost excludes the LLM calls that the agent’s searches trigger in our shared evidence-snippet pipeline (those are measured per run and subtracted; they were 15% to 77% of the billed total, most for the cheapest models) and excludes infrastructure (search, hosting). The thick line is the Pareto frontier: its points are drawn larger, and no other configuration is both cheaper and more accurate than any of them. The line steps because dominance does: a configuration in the red region has some frontier configuration that is at least as cheap and at least as accurate, while anything landing in the green region would join the frontier. FutureSearch SOTA synthesizes forecasts from multiple agent runs, so it has no single-run cost and is not shown.
Pairwise comparisons
Paired bootstrap on pooled scores (numeric weighted 3×)
Each cell is the difference in pooled score (row − column) on the questions both agents forecast; negative (green) means the row agent is more accurate. Bold, bordered cells are statistically significant (two-sided paired-bootstrap * p<.05, ** p<.01, *** p<.001); grey cells are not. Hover a cell for the 95% confidence interval, p-value, and shared question count.
Bench to the Future 2 (BTF-2)
BTF-2 evaluates agents on 1,417 hard forecasting questions. Agents research and forecast offline against a frozen 15M-document corpus. Rationales and reasoning traces are evaluated for strategic reasoning.
BTF-2 Leaderboard
Last updated: 2026-04-20
Agent | Brier (accuracy) | Calibration | Refinement |
|---|---|---|---|
| FutureSearch Agent | 0.119 | 0.002 | 0.081 |
| Opus 4.6 Agent | 0.130 | 0.005 | 0.075 |
| Gemini 3.1 Pro Agent | 0.141 | 0.012 | 0.069 |
| GPT-5.4 Agent | 0.152 | 0.010 | 0.056 |
| Grok 4.20 Beta Agent | 0.165 | 0.003 | 0.039 |
Brier scores on 1,417 pastcasting questions (lower is better). The FutureSearch Agent is an ensemble significantly more accurate than any single frontier agent. Radar chart shows CHAMPS KNOW strategic emphasis (Borda scores, 8 of 10 dimensions).
Papers
Datasets
Deep Research Bench (DRB)
DRB benchmarks how well LLM agents do research on the web. Each of the 169 diverse, real-world tasks provides 10-100k webpages stored offline for search and reasoning, accompanied by carefully curated answers.
DRB Leaderboard
Last updated: 2026-06-10
Agent | Score | Cost ($) | Runtime (s) |
|---|---|---|---|
| Opus 4.6 (high) | 0.553 | $0.53 | 183 |
| Sonnet 4.6 (high) | 0.549 | $0.46 | 262 |
| Opus 4.5 (high) | 0.548 | $0.46 | 140 |
| GPT-5.5 (high) | 0.540 | $0.36 | 231 |
| Opus 4.5 (low) | 0.537 | $0.43 | 162 |
| Opus 4.6 (medium) | 0.532 | $0.36 | 119 |
| Opus 4.6 (low) | 0.514 | $0.24 | 73 |
| Sonnet 4.6 (low) | 0.504 | $0.27 | 130 |
| Opus 4.8 (high) | 0.502 | $0.35 | 121 |
| Gemini 3 Flash (low) | 0.498 | $0.10 | 96 |
| GPT-5 (low) | 0.496 | $0.25 | 230 |
| GPT-5.5 (medium) | 0.496 | $0.31 | 150 |
| Opus 4.8 (low) | 0.493 | $0.24 | 118 |
| Gemini 3 Flash (minimal) | 0.490 | $0.10 | 116 |
| GPT-5 (minimal) | 0.489 | $0.24 | 231 |
| GPT-5.5 (low) | 0.487 | $0.23 | 108 |
| GPT-5 (medium) | 0.486 | $0.35 | 183 |
| Opus 4.1 (medium) | 0.483 | $1.20 | 282 |
| GPT-5 (high) | 0.481 | $0.39 | 217 |
| Gemini 3 Flash (high) | 0.479 | $0.14 | 182 |
| Gemini 3.1 Pro (high) | 0.478 | $0.23 | 187 |
| Sonnet 4.5 (low) | 0.475 | $0.36 | 185 |
| Opus 4.8 (medium) | 0.474 | $0.29 | 106 |
| Grok 4 | 0.473 | $0.49 | 576 |
| Opus 4 (medium) | 0.468 | $1.64 | 283 |
| Sonnet 4 (medium) | 0.466 | $0.45 | 250 |
| Gemini 3 Pro (low) | 0.463 | — | 276 |
| Haiku 4.5 (low) | 0.455 | $0.10 | 83 |
| o3 | 0.452 | $0.12 | 173 |
| Gemini 3.1 Pro (low) | 0.445 | $0.09 | 96 |
| GPT-5.1 (low) | 0.428 | $0.04 | 73 |
| Gemini 2.5 Pro (dynamic) | 0.415 | $0.11 | 191 |
| GPT-5.2 (low) | 0.411 | $0.26 | 136 |
| Gemini 3.1 Flash Lite (low) | 0.373 | $0.02 | 131 |
| Gemini 3.1 Flash Lite (dynamic) | 0.364 | $0.03 | 797 |
| GPT-5.4 Mini (low) | 0.363 | $0.04 | 88 |
| GPT-5.4 (low) | 0.351 | $0.14 | 139 |
Scores averaged first per task category (radar chart), then across all tasks (table). Runtime is estimated from ReAct steps, not wall-clock time.
Papers
Metaculus AI Forecasting Tournaments
Metaculus runs live tournaments where forecasters predict real, unresolved questions and are scored against the field by peer score. Most are bot-only; the Metaculus Cup and the Market Pulse Challenge put bots up against human forecasters, and on those we report where we place against the humans. Our standing in the tournaments we take part in, refreshed at each deploy:
| Tournament | Our standing | Leader |
|---|---|---|
| Summer 2026 FutureEval Bot Tournamentlive $50,000 Prize Pool | #2 of 284 | laertes (5814.64), FutureSearch (5795.03) |
| Metaculus Cup Summer 2026 | Ranked above #3 human | laertes (1098.53), FutureSearch (1021.50) |
| Market Pulse Challenge 26Q3live | Ranked above #2 human | MarcosO (1097.65), FutureSearch (1078.03) |
| MiniBench - 2026-08-24 | #4 of 180 | RMrPeanutButter (1127.46), FutureSearch (948.23) |
| MiniBench - 2026-08-10 | #10 of 178 | allomancer-seon-bot (805.43), FutureSearch (535.86) |
| MiniBench - 2026-07-27 | #6 of 162 | acm_bot (1049.44), FutureSearch (812.56) |
| MiniBench - 2026-07-13 | #2 of 147 | metac-azimuth (831.91), FutureSearch (801.32) |
| MiniBench - 2026-06-29 | #1 of 132 | FutureSearch (732.76) |
| MiniBench - 2026-06-15 | #1 of 116 | FutureSearch (1267.95) |
| MiniBench - 2026-06-01 | #2 of 111 | laertes (740.66), FutureSearch (578.48) |
Standings are pulled from the Metaculus API at deploy time. Entrants are scored by peer score (a per-question comparison against every other forecaster on the same question); higher is better. The leader column shows the top-scoring competitor in the whole field, bot or human; where that isn't us, our score follows for comparison. MiniBench tournaments run on a rolling two-week cadence.
On the Metaculus Cup and Market Pulse, bots and humans forecast the same questions, but only the humans are prize-eligible. “Ranked above #N human” means our score beats the Nth-placed human forecaster in that tournament.
ForecastBench
ForecastBench is a dynamic, contamination-free benchmark of AI forecasting accuracy run by the Forecasting Research Institute. Bots forecast hundreds of unresolved real-world questions, scored on a Brier Index (0–100, higher is better). Our top bot's standing, refreshed at each deploy:
| Leaderboard | Our standing | Leader |
|---|---|---|
| Tournament leaderboard | #18 of 334 | fire hedgehog (69.1), FutureSearch (67.5) |
The tournament leaderboard ranks every submitted entry by its overall Brier Index across dataset and market questions; it is regenerated nightly from the public dataset repository. The Brier Index rescales the mean Brier score to a 0–100 scale (100 = perfect, 50 = uninformed) and adjusts for question difficulty. Our standing is our best-ranked bot; the leader column shows the top-ranked entry, with our score alongside for comparison.
RetroSearch
DRB and BTF-2 use RetroSearch, a system designed to serve agents a frozen, previously scraped version of the internet instead of the live pages, allowing reproducible runs even as the internet changes, and enabling forecasting tasks to be run as "pastcasting".
RetroSearch aims to emulate Google search (specifically, the Serper search API) as closely as possible, so as to minimize differences between live and "retro" agent runs. A single RetroSearch search query follows the following steps:
- Run a live Serper search for the query
- Look up pages obtained from live search in the RetroSearch database and other archive sources
- If the page is not found in the RetroSearch database, remove it from the results
- Write new snippets from a sample of page content using a simple LLM
- Return the results in the original format of the Google results
This approach ensures a search experience for agents that is consistent with real search, but backed exclusively by pages we have a frozen candidate for. The following diagram from the paper illustrates the process:
