---
title: Reinforcement Learning for Optimal Execution
url: https://www.ml-quant.com/papers/ssrn/4720833/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: SSRN 4720833
source_url: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4720833
featured: 2024-02-14
citations: unknown
topic: Trading, Microstructure & Execution
---


# Reinforcement Learning for Optimal Execution

A new actor-critic reinforcement learning algorithm is introduced for optimal execution problem, featuring a recalibration step for convergence and showing linear convergence under appropriate conditions.

- Source: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4720833
- Identifier: SSRN 4720833
- Released: 2023-03-10
- First featured: Quant Letter No. 37 (2024-02-14): https://www.ml-quant.com/issues/2024-02-14/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: Trading, Microstructure & Execution

## Related

- [Reinforcement Learning for Optimal Execution When Liquidity Is Time-Varying](https://www.ml-quant.com/papers/arxiv/2402.12049/): Research shows Double Deep Q-learning, a Reinforcement Learning technique, can effectively learn optimal trading strategies in fluctuating liquidity conditions.
- [The QLBS Model Within the Presence of Feedback Loops Through the Impacts of a Large Trader](https://www.ml-quant.com/papers/arxiv/2311.06790/): The QLBS model is expanded to include a large trader's impact on exchange rates and contingent claim prices, using reinforcement learning to find an optimal hedging strategy, reducing transaction costs and aligning with the trader's fair price.
- [Deviations from the Nash equilibrium in a two-player optimal execution game with reinforcement learning](https://www.ml-quant.com/papers/arxiv/2408.11773/): Autonomous trading bots using advanced algorithms can disrupt markets by deviating from traditional predictions, often favoring optimal solutions over equilibrium.
- [Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2609.22785/): The research extends adversarial reinforcement learning for market making to handle self-exciting order arrivals and price impact, using an LSTM module to improve robustness in complex microstructure environments.
- [Deep Reinforcement Learning for Active High Frequency Trading](https://www.ml-quant.com/papers/arxiv/2101.07107/): A new Deep Reinforcement Learning framework has been developed for high frequency stock trading, showing potential for profitable long-term strategies.
- [Limit Order Book Simulations: A Review](https://www.ml-quant.com/papers/ssrn/4745587/): The piece reviews models of Limit Order Books simulations, emphasizing the role of AI in improving these models and the significance of price impacts in algorithmic trading.
