---
title: The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
url: https://www.ml-quant.com/papers/arxiv/2609.30705/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.30705
source_url: https://arxiv.org/abs/2609.30705
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study tests whether extra reasoning in large language models improves portfolio returns net of trading costs across multiple model families, finding no reliable gains.

- Source: https://arxiv.org/abs/2609.30705
- Identifier: arXiv:2609.30705
- Released: 2026-09-28
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Jiayi Chen, Guiling Wang

## Abstract (arXiv, CC0)

While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.

## Related

- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
- [Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia](https://www.ml-quant.com/papers/arxiv/2312.03664/): Concordia is a library designed to help build and operate Generative Agent-Based Models (GABMs), using Large Language Models (LLMs) to simulate physical or digital environments.
- [Jamba-1.5: Hybrid Transformer-Mamba Models at Scale](https://www.ml-quant.com/papers/arxiv/2408.12570/): Transformer-Mamba Models: Jamba-1.5 is a new large language model with enhanced conversational and instruction-following capabilities, featuring a unique quantization technique for cost-effective inference.
- [MindSearch: Mimicking Human Minds Elicits Deep AI Searcher](https://www.ml-quant.com/papers/arxiv/2407.20183/): Mimicking Human Minds for Search: MindSearch is a Large Language Model-based framework that simulates human cognitive processes for web information seeking, greatly enhancing response quality.
- [Adaptive Partitioning and Learning for Stochastic Control of Diffusion Processes](https://www.ml-quant.com/papers/arxiv/2512.14991/): An adaptive reinforcement learning algorithm enhances learning in controlled diffusion processes by partitioning state-action spaces, providing theoretical guarantees and effective results in applications like portfolio selection.
