arXivML & AI Methods
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
The study tests whether extra reasoning in large language models improves portfolio returns net of trading costs across multiple model families, finding no reliable gains.
Featured in No. 133 on 2 Oct 2026 · 4 days after release

- Released
- 28 Sep 2026
- First featured
- No. 133 · 2 Oct 2026
- Published in
- Not yet, as far as Semantic Scholar knows
- Fanfare
- 3 of 5
- Identifier
- arXiv:2609.30705
- Authors
- Jiayi Chen and Guiling Wang
Abstract
From arXiv (CC0).
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).