ML-QuantSubscribe

arXivML & AI Methods

Can Language Models Learn to Forecast Stock Prices

Post-training Qwen3-4B via supervised fine-tuning and policy optimization more than doubles direction-magnitude score on chronological stock-price predictions to 43.31.

Featured in No. 133 on 2 Oct 2026 · 2 days after release

Post-training brings a 4B model to performance comparable to frontier models on BETA. (a) Scores on the 240 scored test
Figure 1: Post-training brings a 4B model to performance comparable to frontier models on BETA. (a) Scores on the 240 scored test tasks improve from 20.94 to 37.94 to 43.31 through Base, SFT, and PPO. (b) Median calls across all 398 test tasks, including submission, change from 5 to 18 to 16. (c) A…
Released
30 Sep 2026
First featured
No. 133 · 2 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
2 of 5
Identifier
arXiv:2609.36914
Authors
Jiacheng Guo et al.

Abstract

From arXiv (CC0).

Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page