ML-QuantSubscribe

arXivML & AI Methods

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study tests whether extra reasoning in large language models improves portfolio returns net of trading costs across multiple model families, finding no reliable gains.

Featured in No. 133 on 2 Oct 2026 · 4 days after release

Primary low versus baseline return effects across 241 return dates
Figure 2 . Primary low versus baseline return effects across 241 return dates. Circles denote numerical inputs, squares denote identifiable news, and diamonds denote masked news. Horizontal bars are two-sided 95% HAC CIs, which account for changing variance and serial dependence. The solid line is…
Released
28 Sep 2026
First featured
No. 133 · 2 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2609.30705
Authors
Jiayi Chen and Guiling Wang

Abstract

From arXiv (CC0).

While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page