---
title: Domain-adapted Learning and Imitation: DRL for Power Arbitrage
url: https://www.ml-quant.com/papers/arxiv/2301.08360/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2301.08360
source_url: https://arxiv.org/abs/2301.08360
featured: 2023-09-14
citations: 1
topic: Trading, Microstructure & Execution
---


# Domain-adapted Learning and Imitation: DRL for Power Arbitrage

Leveraging Expertise: A dual-agent reinforcement learning approach can optimize European power arbitrage trading, improving training convergence and performance, and tripling profit and loss.

- Source: https://arxiv.org/abs/2301.08360
- Identifier: arXiv:2301.08360
- Released: 2023-01-19
- First featured: Quant Letter No. 15 (2023-09-14): https://www.ml-quant.com/issues/2023-09-14/
- Citations (Semantic Scholar): 1
- Published in: not yet
- Topic: Trading, Microstructure & Execution

## Related

- [Deep Reinforcement Learning for Active High Frequency Trading](https://www.ml-quant.com/papers/arxiv/2101.07107/): A new Deep Reinforcement Learning framework has been developed for high frequency stock trading, showing potential for profitable long-term strategies.
- [Gray-box Adversarial Attack of Deep Reinforcement Learning-based Trading Agents*](https://www.ml-quant.com/papers/arxiv/2309.14615/): A study has shown that a gray-box method can significantly reduce the profits of a Deep Reinforcement Learning-based trading agent, highlighting the need for stronger automated trading systems.
- [JAX-LOB: A GPU-Accelerated limit order book simulator to unlock large scale reinforcement learning for trading](https://www.ml-quant.com/papers/arxiv/2308.13289/): JAX-LOB: The paper introduces JAX-LOB, the first GPU-powered limit order book simulator capable of processing multiple books simultaneously, designed for efficient large-scale simulations of LOB dynamics for research, calibration, and reinforcement learning training.
- [An adaptive dual-level reinforcement learning approach for optimal trade execution](https://www.ml-quant.com/papers/arxiv/2307.10649/): The research introduces a reinforcement learning strategy that accurately tracks the daily volume-weighted average price of stocks, using a dual-level architecture for better results.
- [The QLBS Model Within the Presence of Feedback Loops Through the Impacts of a Large Trader](https://www.ml-quant.com/papers/arxiv/2311.06790/): The QLBS model is expanded to include a large trader's impact on exchange rates and contingent claim prices, using reinforcement learning to find an optimal hedging strategy, reducing transaction costs and aligning with the trader's fair price.
- [Deep Reinforcement Learning: Policy Gradients for US Equities Trading](https://www.ml-quant.com/papers/ssrn/4645453/): The study shows that Deep Reinforcement Learning can effectively interpret synthetic alpha signals in financial trading, outperforming the market benchmark.
