---
title: How Execution Assumptions Change Short-Horizon Sharpe Rankings: Evidence from a Synthetic Trading Benchmark
url: https://www.ml-quant.com/papers/arxiv/2610.05077/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-09
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2610.05077
source_url: https://arxiv.org/abs/2610.05077
featured: 2026-10-09
citations: unknown
topic: Trading, Microstructure & Execution
---


# How Execution Assumptions Change Short-Horizon Sharpe Rankings: Evidence from a Synthetic Trading Benchmark

Varying execution realism from ideal fills to latency and impact reshuffles rankings of LLM and classical trading policies, showing how backtest conventions affect headline results.

- Source: https://arxiv.org/abs/2610.05077
- Identifier: arXiv:2610.05077
- Released: 2026-10-06
- First featured: Quant Letter No. 134 (2026-10-09): https://www.ml-quant.com/issues/2026-10-09/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: Trading, Microstructure & Execution
- Authors: Weicheng Xue

## Abstract (arXiv, CC0)

Backtests of LLM trading agents often assume that every order fills at the closing price. We ask whether this choice changes only reported returns or also the order of the agents. Five prompted LLM signal policies and seven classical baselines trade the same synthetic price paths under six execution settings, from near-ideal fills to latency, spread, participation, and impact stresses. The main experiment contains $2{,}462$ runs with matched decision frequencies and paired market paths. On the compressed two-asset board, agreement between the near-ideal and default-stress rankings falls to Kendall $τ_b=0.21$ in the high-volatility regime, compared with $0.82$ in the calm regime. The seed-bootstrap intervals, $[0.00,0.52]$ and $[0.48,0.94]$, are wide and overlap. On a fixed 11-policy board, agreement rises from 0.24 with two assets to 0.85 with ten; the two-asset point estimate differs substantially from the wider settings we tested. Rank changes are related to turnover, and comparisons with buy-and-hold also depend on how that anchor is initialized. The experiment does not compare LLM trading skill. It shows that, on a short horizon, an execution convention can become part of the benchmark's headline. Execution assumptions and rank stability should be reported alongside returns.

## Related

- [ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism](https://www.ml-quant.com/papers/arxiv/2508.00554/): The article introduces a new multi-agent system for financial trading that uses competition to analyze market data and make decisions, proving more effective than traditional methods in tests.
- [Classifying and clustering trading agents](https://www.ml-quant.com/papers/arxiv/2505.21662/): The research suggests using an agent-based model to create synthetic data for categorizing financial investors by their behavior, emphasizing the difficulties in validating and interpreting machine learning methods.
- [The bias of IID resampled backtests for rolling window mean-variance portfolios](https://www.ml-quant.com/papers/arxiv/2505.06383/): A study finds that resampling techniques in backtests can cause bias in Sharpe Ratio estimates, suggesting a need for structure-preserving resampling methods.
- [Optimizing Portfolio Performance through Clustering and Sharpe Ratio-Based Optimization: A Comparative Backtesting Approach](https://www.ml-quant.com/papers/arxiv/2501.12074/): The article presents a new method for improving portfolio performance using clustering-based segmentation and Sharpe ratio-based optimization, tested with historical data from various asset classes.
- [Tax Credits and Household Behavior: The Roles of Myopic Decision-Making and Liquidity in a Simulated Economy](https://www.ml-quant.com/papers/arxiv/2408.10391/): The research uses a simulator with AI agents to study the effects of tax credits on various households, suggesting a new distribution method to lessen inequality.
- [Forecasting and backtesting gradient allocations of expected shortfall](https://www.ml-quant.com/papers/arxiv/2401.11701/): The study introduces a new semiparametric model for predicting dynamic Expected Shortfall contributions, which has shown excellent results in analyzing stock returns.
