---
title: Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
url: https://www.ml-quant.com/papers/arxiv/2609.28876/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.28876
source_url: https://arxiv.org/abs/2609.28876
featured: 2026-09-25
citations: 0
topic: ML & AI Methods
---


# Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

Introduces a replayable environment combining 1,568 resolved prediction-market questions with 18.8M dated news articles to benchmark and train language-model forecasting agents on historical data.

- Source: https://arxiv.org/abs/2609.28876
- Identifier: arXiv:2609.28876
- Released: 2026-09-25
- First featured: Quant Letter No. 132 (2026-09-25): https://www.ml-quant.com/issues/2026-09-25/
- Citations (Semantic Scholar): 0
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Liqin Ye, Haorui Wang, Fardin Ahmed, Rongzhi Zhang, Yuan He, Ziyuan Lin, Yanbin Yin, Jing Peng, Michael Galarnyk, Sudheer Chava, Chao Zhang

## Abstract (arXiv, CC0)

We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.

## Related

- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
- [Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia](https://www.ml-quant.com/papers/arxiv/2312.03664/): Concordia is a library designed to help build and operate Generative Agent-Based Models (GABMs), using Large Language Models (LLMs) to simulate physical or digital environments.
- [Jamba-1.5: Hybrid Transformer-Mamba Models at Scale](https://www.ml-quant.com/papers/arxiv/2408.12570/): Transformer-Mamba Models: Jamba-1.5 is a new large language model with enhanced conversational and instruction-following capabilities, featuring a unique quantization technique for cost-effective inference.
- [MindSearch: Mimicking Human Minds Elicits Deep AI Searcher](https://www.ml-quant.com/papers/arxiv/2407.20183/): Mimicking Human Minds for Search: MindSearch is a Large Language Model-based framework that simulates human cognitive processes for web information seeking, greatly enhancing response quality.
- [From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets](https://www.ml-quant.com/papers/arxiv/2509.06069/): Large language models (LLMs) in expert services have pros and cons, with human markets being more efficient, but LLMs potentially reducing trust and overshadowing experts' preferences, while also improving experts' communication of their goals.
