---
title: AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
url: https://www.ml-quant.com/papers/arxiv/2609.34188/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.34188
source_url: https://arxiv.org/abs/2609.34188
featured: 2026-10-02
citations: unknown
topic: ML & AI Methods
---


# AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

The research uses reinforcement learning with multi-objective rewards to discover formulaic alphas that work well together despite evolving reward functions and shifting environments.

- Source: https://arxiv.org/abs/2609.34188
- Identifier: arXiv:2609.34188
- Released: 2026-09-29
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Yingbo Zhao, Zeyu Yang, Zhoufan Zhu

## Abstract (arXiv, CC0)

Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.

## Related

- [VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback](https://www.ml-quant.com/papers/arxiv/2409.18417/): An auction mechanism is introduced to enhance cost-efficiency in fine-tuning large language models using Reinforcement Learning from Human Feedback, focusing on quality feedback and model performance.
- [MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions](https://www.ml-quant.com/papers/arxiv/2410.02743/): The MA-RLHF framework integrates macro actions into the learning process of large language models, enhancing learning efficiency and performance in tasks like text summarization and dialogue generation.
- [Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets](https://www.ml-quant.com/papers/arxiv/2609.34510/): A benchmark compares machine learning, reinforcement learning, and large language model trading methods across historical backtests, paper trading, and live markets to measure the gap.
- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
