ML-QuantSubscribe

arXivTrading, Microstructure & Execution

How Execution Assumptions Change Short-Horizon Sharpe Rankings: Evidence from a Synthetic Trading Benchmark

Varying execution realism from ideal fills to latency and impact reshuffles rankings of LLM and classical trading policies, showing how backtest conventions affect headline results.

Featured in No. 134 on 9 Oct 2026 · 3 days after release

Kendall \tau_{b} between the E0 ranking and each stressed ranking
Figure 1. Kendall \tau_{b} between the E0 ranking and each stressed ranking. Rows are market regimes and columns are execution settings. Each cell contains 12 policies ranked by mean Sharpe over ten seeds; red cells indicate less agreement and green cells more. A three-by-five heatmap of Kendall ta…
Released
6 Oct 2026
First featured
No. 134 · 9 Oct 2026
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2610.05077
Authors
Weicheng Xue

Abstract

From arXiv (CC0).

Backtests of LLM trading agents often assume that every order fills at the closing price. We ask whether this choice changes only reported returns or also the order of the agents. Five prompted LLM signal policies and seven classical baselines trade the same synthetic price paths under six execution settings, from near-ideal fills to latency, spread, participation, and impact stresses. The main experiment contains $2{,}462$ runs with matched decision frequencies and paired market paths. On the compressed two-asset board, agreement between the near-ideal and default-stress rankings falls to Kendall $τ_b=0.21$ in the high-volatility regime, compared with $0.82$ in the calm regime. The seed-bootstrap intervals, $[0.00,0.52]$ and $[0.48,0.94]$, are wide and overlap. On a fixed 11-policy board, agreement rises from 0.24 with two assets to 0.85 with ten; the two-asset point estimate differs substantially from the wider settings we tested. Rank changes are related to turnover, and comparisons with buy-and-hold also depend on how that anchor is initialized. The experiment does not compare LLM trading skill. It shows that, on a short horizon, an execution convention can become part of the benchmark's headline. Execution assumptions and rank stability should be reported alongside returns.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page