---
title: LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs
url: https://www.ml-quant.com/papers/arxiv/2609.33470/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.33470
source_url: https://arxiv.org/abs/2609.33470
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

An evaluation framework tests language model agents on structured option trading tasks, finding current systems underperform in most real-world scenarios.

- Source: https://arxiv.org/abs/2609.33470
- Identifier: arXiv:2609.33470
- Released: 2026-09-29
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Haochen Luo, Yifan Li, Binh Minh An, Xiaolong Luo, Zhengzhao Lai, Yuan Zhang, Chen Liu

## Abstract (arXiv, CC0)

Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challenges, including nonlinear payoffs and multi-leg strategy construction, requiring structured decisions rather than simple directional bets. We introduce LiveOption, an evaluation framework for LLM-based agents in option trading. LiveOption formulates the problem as structured sequential decision-making under realistic execution and capital constraints, and provides a reproducible environment with standardized interaction protocols. The framework includes three task suites covering portfolio overlays, event-driven earnings trading, and 0DTE intraday trading. We further propose a hierarchical metric suite that evaluates action validity, decision quality, risk characteristics, and outcome-level performance. Experiments show that current agents often fail to achieve competitive returns in most scenarios. LiveOption offers a principled testbed for evaluating structured decision-making beyond outcome-based metrics.

## Related

- [Can Generative AI agents behave like humans? Evidence from laboratory market experiments](https://www.ml-quant.com/papers/arxiv/2505.07457/): Large Language Models (LLMs) have potential in mimicking human behavior in economic markets, but need more research for improved diversity and accuracy.
- [Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets](https://www.ml-quant.com/papers/arxiv/2609.34510/): A benchmark compares machine learning, reinforcement learning, and large language model trading methods across historical backtests, paper trading, and live markets to measure the gap.
- [Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors](https://www.ml-quant.com/papers/arxiv/2609.27051/): Proposes a statistical referee that judges investment factors proposed by language-model agents using out-of-sample market outcomes, ensuring false-discovery control at any stopping time.
- [AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining](https://www.ml-quant.com/papers/arxiv/2609.29014/): Proposes a multi-agent system with post-training that automates alpha factor mining locally, using diverse research paths and joint optimization to broaden exploration while maintaining prediction quality.
- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
