---
title: Multiperiod bond portfolio optimization with transaction costs using a Markov Decision process
url: https://www.ml-quant.com/papers/arxiv/2609.38765/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-10-02
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2609.38765
source_url: https://arxiv.org/abs/2609.38765
featured: 2026-10-02
citations: unknown
topic: Portfolio & Allocation
---


# Multiperiod bond portfolio optimization with transaction costs using a Markov Decision process

The study develops a Markov decision process framework for dynamic bond portfolio management that balances yield and interest-rate risk subject to transaction costs.

- Source: https://arxiv.org/abs/2609.38765
- Identifier: arXiv:2609.38765
- Released: 2026-10-01
- First featured: Quant Letter No. 133 (2026-10-02): https://www.ml-quant.com/issues/2026-10-02/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: Portfolio & Allocation
- Authors: Balaji Ramachandran, Srikanth Iyer, Shashi Jain

## Abstract (arXiv, CC0)

Bank treasury portfolios must balance yield, liquidity, and interest-rate risk across bonds of different maturities. Static allocation rules are ill-suited to this task: portfolios concentrated in long-duration securities with no dynamic adjust- ment mechanism can accumulate large mark-to-market losses and liquidity stress under rising interest rates, as illustrated by the failure of Silicon Valley Bank in 2023. We develop a tractable simulation-based framework for multi-period bond port- folio optimization under interest-rate risk and proportional transaction costs. Yield-curve dynamics are modeled using the Dynamic Nelson-Siegel parameter- ization with Vector Autoregressive factor dynamics, from which we construct a time-inhomogeneous discrete-state Markov chain approximating the joint yield process across bond maturities. This chain forms the state space of a finite- horizon Markov Decision Process in which the investor maximizes expected terminal wealth subject to proportional rebalancing costs. The optimal portfolio policy is obtained by backward induction. We also quantify the approximation error introduced by truncating the transition kernel, and show that it leaves mean terminal wealth almost unchanged while substantially distorting drawdown and tail statistics.

## Related

- [Portfolio Optimization under Transaction Costs with Recursive Preferences](https://www.ml-quant.com/papers/arxiv/2402.08387/): The Merton investment-consumption problem is expanded to incorporate transaction costs and stochastic differential utility, using new math techniques to understand all parameter combinations and previously difficult aspects.
- [Portfolio Construction: Low Risk High Variability](https://www.ml-quant.com/papers/ssrn/5105457/): Low Risk High Variability: Research indicates that stocks with less volatility yield higher returns, with portfolio construction and transaction costs significantly impacting low-risk portfolio performance.
- [Adaptive Online Portfolio Selection](https://www.ml-quant.com/papers/repec/eee-ejores-v-321-y-2025-i-1-p-214-230/): The article introduces a new online portfolio selection strategy that considers transaction costs and uses an adaptive scheme for sequential parameter decision, yielding higher cumulative returns and competitive Sharpe ratios than existing strategies.
- [Adaptive Online Portfolio Selection with Transaction Costs](https://www.ml-quant.com/papers/repec/taf-quantf-v-24-y-2023-i-1-p-59-82/): The research proposes a new algorithm for online portfolio selection that improves return prediction accuracy by considering peer impact.
- [Optimizing Investment Strategies with Lazy Factor and Probability Weighting: A Price Portfolio Forecasting and Mean-Variance Model with Transaction Costs Approach](https://www.ml-quant.com/papers/arxiv/2306.07928/): A new investment strategy model has been developed and tested on a dataset.
- [The Critical Line Algorithm and the Constrained LASSO: One Curve, Two Literatures](https://www.ml-quant.com/papers/arxiv/2609.25704/): Shows that mean-variance portfolio selection and the constrained LASSO trace identical piecewise-linear solution paths, mapping their parametrizations exactly.
