---
title: PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading
url: https://www.ml-quant.com/papers/arxiv/2610.01325/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2610.01325
source_url: https://arxiv.org/abs/2610.01325
featured: 2026-10-02
citations: unknown
topic: Trading, Microstructure & Execution
---


# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading

Reinforcement learning agent combining policy optimization with regime priors controls drawdown while blending learned and rule-based portfolio exposure on held-out equity data.

- Source: https://arxiv.org/abs/2610.01325
- Identifier: arXiv:2610.01325
- Released: 2026-10-02
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: Trading, Microstructure & Execution
- Authors: Duong Hien Chi Kien, Thanh Trung Huynh

## Abstract (arXiv, CC0)

Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.

## Related

- [Domain-adapted Learning and Interpretability: DRL for Gas Trading](https://www.ml-quant.com/papers/arxiv/2301.08359/): Enhanced Performance: Deep Reinforcement Learning (Deep RL) can enhance trading of natural gas futures contracts, outperforming traditional strategies through ensemble learning.
- [Deep Reinforcement Learning for Active High Frequency Trading](https://www.ml-quant.com/papers/arxiv/2101.07107/): A new Deep Reinforcement Learning framework has been developed for high frequency stock trading, showing potential for profitable long-term strategies.
- [Reinforcement learning for trade execution with market and limit orders](https://www.ml-quant.com/papers/arxiv/2507.06345/): The paper presents a reinforcement learning framework for optimal trade execution, using multivariate logistic-normal distributions, which outperforms traditional benchmark strategies.
- [Signature Decomposition Method Applying to Pair Trading](https://www.ml-quant.com/papers/arxiv/2505.05332/): A new pairs trading strategy using path signature techniques enhances futures trading by providing better interpretability, robustness, and returns.
- [Learning the Spoofability of Limit Order Books With Interpretable Probabilistic Neural Networks](https://www.ml-quant.com/papers/arxiv/2504.15908/): The article discusses a new real-time detection model for spotting potential market manipulation in cryptocurrency exchanges, with 31% of large orders identified as potential spoofs.
- [Deviations from the Nash equilibrium in a two-player optimal execution game with reinforcement learning](https://www.ml-quant.com/papers/arxiv/2408.11773/): Autonomous trading bots using advanced algorithms can disrupt markets by deviating from traditional predictions, often favoring optimal solutions over equilibrium.
