---
title: Consistent time travel for realistic interactions with historical data: reinforcement learning for market making
url: https://www.ml-quant.com/papers/arxiv/2408.02322/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2408.02322
source_url: https://arxiv.org/abs/2408.02322
featured: 2024-08-07
citations: 0
topic: Trading, Microstructure & Execution
---


# Consistent time travel for realistic interactions with historical data: reinforcement learning for market making

The article discusses the use of consistent data time travel in offline reinforcement learning for market making in limit order books.

- Source: https://arxiv.org/abs/2408.02322
- Identifier: arXiv:2408.02322
- Released: 2024-08-05
- First featured: Quant Letter No. 60 (2024-08-07): https://www.ml-quant.com/issues/2024-08-07/
- Citations (Semantic Scholar): 0
- Published in: not yet
- Topic: Trading, Microstructure & Execution

## Related

- [Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2609.22785/): The research extends adversarial reinforcement learning for market making to handle self-exciting order arrivals and price impact, using an LSTM module to improve robustness in complex microstructure environments.
- [Market Making with Exogenous Competition](https://www.ml-quant.com/papers/arxiv/2407.17393/): A study reveals that a 'reference market maker' who optimizes her posted depths can achieve a near perfect solution, which is compared against other solutions using an Euler scheme or reinforcement learning techniques in a competitive environment.
- [The Negative Drift of a Limit Order Fill](https://www.ml-quant.com/papers/arxiv/2407.16527/): The research identifies a negative drift in market making models, particularly in limit order fills, using the 10 Year US Treasury Bond futures for empirical simulation.
- [Event-Based Limit Order Book Simulation under a Neural Hawkes Process: Application in Market-Making](https://www.ml-quant.com/papers/arxiv/2502.17417/): An event-driven Limit Order Book model using a Neural Hawkes process is proposed to simulate high-frequency dynamics in financial markets, offering a more accurate depiction of trade execution.
- [Reinforcement learning for trade execution with market and limit orders](https://www.ml-quant.com/papers/arxiv/2507.06345/): The paper presents a reinforcement learning framework for optimal trade execution, using multivariate logistic-normal distributions, which outperforms traditional benchmark strategies.
- [JAX-LOB: A GPU-Accelerated limit order book simulator to unlock large scale reinforcement learning for trading](https://www.ml-quant.com/papers/arxiv/2308.13289/): JAX-LOB: The paper introduces JAX-LOB, the first GPU-powered limit order book simulator capable of processing multiple books simultaneously, designed for efficient large-scale simulations of LOB dynamics for research, calibration, and reinforcement learning training.
