{"schemaVersion":"2026-07-21.signal.v2","id":"live-e8bdffd2d44d17e9e939","title":"Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents","slug":"arxiv-cs-ai-oai-arxiv-org-2609-28876v1-forecast-dojo-replayable-environments-for-bench-e9e939","url":"https://www.niubiagent.com/signals/arxiv-cs-ai-oai-arxiv-org-2609-28876v1-forecast-dojo-replayable-environments-for-bench-e9e939","jsonUrl":"https://www.niubiagent.com/api/posts/arxiv-cs-ai-oai-arxiv-org-2609-28876v1-forecast-dojo-replayable-environments-for-bench-e9e939.json","markdownUrl":"https://www.niubiagent.com/content/arxiv-cs-ai-oai-arxiv-org-2609-28876v1-forecast-dojo-replayable-environments-for-bench-e9e939","summaryHuman":"arXiv Computer Science AI published Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents. arXiv:2609.28876v1 Announce Type: new Abstract: We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing…","summaryAgent":"Treat Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents as an official publication signal. Read the primary source, verify the announced change, and assess whether it affects your agent stack.","category":"safety-research","tags":["arxiv","research","agents"],"sourceName":"arXiv Computer Science AI","sourceUrl":"https://arxiv.org/abs/2609.28876","publishedAt":"2026-09-25T04:00:00.000Z","curatedAt":"2026-09-26T00:17:52.116Z","confidence":0.9,"agentUsefulness":55,"sponsorIds":[],"language":"en","contentMode":"source-watch","verifiedAt":"2026-09-26T00:17:52.116Z","changeType":"ecosystem","actionItems":["Read the original arXiv Computer Science AI article before relying on this summary.","Verify the announced capabilities and dates against the primary source.","Assess whether the change affects your agent stack or evaluation plan."],"body":"arXiv Computer Science AI published Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents. This automated source-watch entry was generated from the publisher's official RSS feed and is not human-reviewed editorial analysis. Source excerpt: arXiv:2609.28876v1 Announce Type: new Abstract: We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and…","sponsors":[]}