---
title: Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
url: https://www.ml-quant.com/papers/arxiv/2609.34510/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.34510
source_url: https://arxiv.org/abs/2609.34510
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets

A benchmark compares machine learning, reinforcement learning, and large language model trading methods across historical backtests, paper trading, and live markets to measure the gap.

- Source: https://arxiv.org/abs/2609.34510
- Identifier: arXiv:2609.34510
- Released: 2026-09-29
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Xingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei, Jiarui Liu, Chang Zhou, Fangzhou Ge, Chenyi Xu, Xikun Zhang, Renqiang Luo, Jie Zhang, Hong Cheng

## Abstract (arXiv, CC0)

AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at https://github.com/Starlien95/Awesome-TradingAI.

## Related

- [Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors](https://www.ml-quant.com/papers/arxiv/2609.27051/): Proposes a statistical referee that judges investment factors proposed by language-model agents using out-of-sample market outcomes, ensuring false-discovery control at any stopping time.
- [Can Generative AI agents behave like humans? Evidence from laboratory market experiments](https://www.ml-quant.com/papers/arxiv/2505.07457/): Large Language Models (LLMs) have potential in mimicking human behavior in economic markets, but need more research for improved diversity and accuracy.
- [VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback](https://www.ml-quant.com/papers/arxiv/2409.18417/): An auction mechanism is introduced to enhance cost-efficiency in fine-tuning large language models using Reinforcement Learning from Human Feedback, focusing on quality feedback and model performance.
- [MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions](https://www.ml-quant.com/papers/arxiv/2410.02743/): The MA-RLHF framework integrates macro actions into the learning process of large language models, enhancing learning efficiency and performance in tasks like text summarization and dialogue generation.
- [AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2609.34188/): The research uses reinforcement learning with multi-objective rewards to discover formulaic alphas that work well together despite evolving reward functions and shifting environments.
- [LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs](https://www.ml-quant.com/papers/arxiv/2609.33470/): An evaluation framework tests language model agents on structured option trading tasks, finding current systems underperform in most real-world scenarios.
