---
title: Markowitz Meets Bellman: Knowledge-distilled Reinforcement Learning for Portfolio Management
url: https://www.ml-quant.com/papers/arxiv/2405.05449/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2405.05449
source_url: https://arxiv.org/abs/2405.05449
featured: 2024-05-15
citations: 1
topic: Portfolio & Allocation
---


# Markowitz Meets Bellman: Knowledge-distilled Reinforcement Learning for Portfolio Management

The paper presents KDD, a hybrid method combining portfolio theory and reinforcement learning for optimal investment portfolios, achieving high profitability with low risk.

- Source: https://arxiv.org/abs/2405.05449
- Identifier: arXiv:2405.05449
- Released: 2024-05-08
- First featured: Quant Letter No. 49 (2024-05-15): https://www.ml-quant.com/issues/2024-05-15/
- Citations (Semantic Scholar): 1
- Published in: not yet
- Topic: Portfolio & Allocation

## Related

- [Deep Reinforcement Learning: Extending Traditional Financial Portfolio Methods](https://www.ml-quant.com/papers/ssrn/4780026/): The paper suggests that deep reinforcement learning can potentially improve traditional portfolio allocation strategies by incorporating contextual data and future rewards.
- [Explainable post hoc portfolio management financial policy of a Deep Reinforcement Learning agent](https://www.ml-quant.com/papers/arxiv/2407.14486/): A new Explainable Deep Reinforcement Learning (XDRL) method for portfolio management has been developed, combining Proximal Policy Optimization with explainable techniques for better transparency in investment predictions.
- [Developing A Multi-Agent and Self-Adaptive Framework with Deep Reinforcement Learning for Dynamic Portfolio Risk Management](https://www.ml-quant.com/papers/arxiv/2402.00515/): The piece introduces a multi-agent and self-adaptive framework (MASA) for portfolio management, using reinforcement learning to balance returns and risks, providing market trend feedback and outperforming other similar approaches.
- [Discrete-Time Mean-Variance Strategy Based on Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2312.15385/): The article discusses a new reinforcement learning-based model for analyzing real-world data, which is more applicable than the continuous-time model.
- [Data-Driven Merton's Strategies via Policy Randomization](https://www.ml-quant.com/papers/arxiv/2312.11797/): The study applies reinforcement learning to determine optimal portfolio policies in an incomplete market, showing its efficiency and robustness compared to the traditional plug-in method.
- [On Optimal Tracking Portfolio in Incomplete Markets: The Reinforcement Learning Approach](https://www.ml-quant.com/papers/arxiv/2311.14318/): The paper uses capital injection to solve an optimal tracking portfolio problem in incomplete market models, showcasing the q-learning algorithm's performance.
