arXivML & AI Methods
Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
Introduces a replayable environment combining 1,568 resolved prediction-market questions with 18.8M dated news articles to benchmark and train language-model forecasting agents on historical data.
Featured in No. 132 on 25 Sep 2026 · on release day · 0 citations today

- Released
- 25 Sep 2026
- First featured
- No. 132 · 25 Sep 2026
- Citations (Semantic Scholar)
- 0
- Influential citations
- 0
- Published in
- Not yet, as far as Semantic Scholar knows
- Fanfare
- 3 of 5
- Identifier
- arXiv:2609.28876
- Authors
- Liqin Ye et al.
Abstract
From arXiv (CC0).
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).