---
title: Does Autonomous Quant Research Improve? Leakage, Evidence Scarcity and Forward Tests in an LLM Factor-Discovery Loop
url: https://www.ml-quant.com/papers/ssrn/7555699/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-09
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: SSRN 7555699
source_url: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7555699
featured: 2026-10-09
citations: unknown
topic: ML & AI Methods
---


# Does Autonomous Quant Research Improve? Leakage, Evidence Scarcity and Forward Tests in an LLM Factor-Discovery Loop

An autonomous agent evolved 940 factors over 17 days and finds that reusing backtest data inflates edge by a quarter to a third and in-sample improvement predicts worse performance.

- Source: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7555699
- Identifier: SSRN 7555699
- Released: 2026-10-05
- First featured: Quant Letter No. 134 (2026-10-09): https://www.ml-quant.com/issues/2026-10-09/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Kamer Ali Yuksel

## Related

- [Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets](https://www.ml-quant.com/papers/arxiv/2609.34510/): A benchmark compares machine learning, reinforcement learning, and large language model trading methods across historical backtests, paper trading, and live markets to measure the gap.
- [Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors](https://www.ml-quant.com/papers/arxiv/2609.27051/): Proposes a statistical referee that judges investment factors proposed by language-model agents using out-of-sample market outcomes, ensuring false-discovery control at any stopping time.
- [Can Generative AI agents behave like humans? Evidence from laboratory market experiments](https://www.ml-quant.com/papers/arxiv/2505.07457/): Large Language Models (LLMs) have potential in mimicking human behavior in economic markets, but need more research for improved diversity and accuracy.
- [LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets](https://www.ml-quant.com/papers/arxiv/2610.09872/): A benchmark evaluates frontier LLMs as live trading agents and finds that realized returns often diverge from capability-specific measurements, revealing outcome-capability gaps through decision traces.
- [Trading Strategy Optimization via Textual Gradient](https://www.ml-quant.com/papers/arxiv/2610.03128/): Proposes an LLM-guided framework for trading strategy refinement using accumulated experience and cross-period robustness, achieving 27.99% annualized return on Chinese equities with 1.63 Sharpe ratio.
- [LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs](https://www.ml-quant.com/papers/arxiv/2609.33470/): An evaluation framework tests language model agents on structured option trading tasks, finding current systems underperform in most real-world scenarios.
