---
title: Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?
url: https://www.ml-quant.com/papers/arxiv/2505.07078/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2505.07078
source_url: http://arxiv.org/abs/2505.07078v1
featured: 2025-05-14
citations: 26
topic: LLMs & Text
---


# Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?

FINSABER, a backtesting framework, shows that Large Language Models' effectiveness in stock trading decreases over longer periods and larger symbol universes, emphasizing the need for trend detection and risk controls.

- Source: http://arxiv.org/abs/2505.07078v1
- Identifier: arXiv:2505.07078
- Released: 2025-05-11
- First featured: Quant Letter No. 97 (2025-05-14): https://www.ml-quant.com/issues/2025-05-14/
- Citations (Semantic Scholar): 26
- Published in: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
- Topic: LLMs & Text

## Related

- [The Memorization Problem: Can We Trust LLMs'Economic Forecasts?](https://www.ml-quant.com/papers/arxiv/2504.14765/): The study shows that large language models can remember exact economic figures from before their knowledge cutoff dates, which may skew their predictive abilities in forecasting and backtesting trading strategies.
- [Look-Ahead Bias in Stock Return Predictions](https://www.ml-quant.com/papers/ssrn/4586726/): Large language models like ChatGPT can generate profitable trading signals from news sentiment, but backtesting can yield biased results due to overlapping periods.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
