---
title: MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
url: https://www.ml-quant.com/papers/arxiv/2610.03080/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-09
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2610.03080
source_url: https://arxiv.org/abs/2610.03080
featured: 2026-10-09
citations: unknown
topic: LLMs & Text
---


# MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

Introduces a benchmark that compares LLM-generated trading code to reference strategies bar by bar on identical data, revealing silent failures in implementation despite passing functional tests.

- Source: https://arxiv.org/abs/2610.03080
- Identifier: arXiv:2610.03080
- Released: 2026-10-05
- First featured: Quant Letter No. 134 (2026-10-09): https://www.ml-quant.com/issues/2026-10-09/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: LLMs & Text
- Authors: Siyu Wang, Yifan Wang, Yuecheng He

## Abstract (arXiv, CC0)

Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.

## Related

- [Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?](https://www.ml-quant.com/papers/arxiv/2505.07078/): FINSABER, a backtesting framework, shows that Large Language Models' effectiveness in stock trading decreases over longer periods and larger symbol universes, emphasizing the need for trend detection and risk controls.
- [The Memorization Problem: Can We Trust LLMs'Economic Forecasts?](https://www.ml-quant.com/papers/arxiv/2504.14765/): The study shows that large language models can remember exact economic figures from before their knowledge cutoff dates, which may skew their predictive abilities in forecasting and backtesting trading strategies.
- [Look-Ahead Bias in Stock Return Predictions](https://www.ml-quant.com/papers/ssrn/4586726/): Large language models like ChatGPT can generate profitable trading signals from news sentiment, but backtesting can yield biased results due to overlapping periods.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
