---
title: Stochastic Control with Exit Time: Policy Gradient Learning
url: https://www.ml-quant.com/papers/repec/taf-apmtfi-v-29-y-2022-i-6-p-439-456/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: RePEc:taf:apmtfi:v:29:y:2022:i:6:p:439-456
source_url: https://econpapers.repec.org/scripts/redir.pf?u=http%3A%2F%2Fhdl.handle.net%2F10.1080%2F1350486X.2023.2239850%3Bh%3Drepec%3Ataf%3Aapmtfi%3Av%3A29%3Ay%3A2022%3Ai%3A6%3Ap%3A439-456
featured: 2023-09-21
citations: unknown
topic: ML & AI Methods
---


# Stochastic Control with Exit Time: Policy Gradient Learning

Policy Gradient Learning: The research shows that policy gradient methods for stochastic control with exit time outperform other techniques in share repurchase pricing and can adapt to realistic market conditions.

- Source: https://econpapers.repec.org/scripts/redir.pf?u=http%3A%2F%2Fhdl.handle.net%2F10.1080%2F1350486X.2023.2239850%3Bh%3Drepec%3Ataf%3Aapmtfi%3Av%3A29%3Ay%3A2022%3Ai%3A6%3Ap%3A439-456
- Identifier: RePEc:taf:apmtfi:v:29:y:2022:i:6:p:439-456
- Released: 2022-12-12
- First featured: Quant Letter No. 16 (2023-09-21): https://www.ml-quant.com/issues/2023-09-21/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods

## Related

- [Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control](https://www.ml-quant.com/papers/arxiv/2409.08861/): The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
- [Optimal Stopping via Randomized Neural Networks](https://www.ml-quant.com/papers/arxiv/2104.13669/): The article highlights the benefits of using randomized neural networks to approximate solutions for optimal stopping problems, proving they are more efficient and faster than other machine learning methods.
- [Deep Penalty Methods: A Class of Deep Learning Algorithms for Solving High Dimensional Optimal Stopping Problems](https://www.ml-quant.com/papers/arxiv/2405.11392/): A proposed deep learning algorithm for optimal stopping problems shows accuracy and efficiency in American option pricing, with its error bound by the loss function and other parameters.
- [Optimal stopping and divestment timing under scenario ambiguity and learning](https://www.ml-quant.com/papers/arxiv/2408.09349/): The research analyzes the effect of environmental transition on asset value, using a decision-making model to determine optimal divestment decisions.
- [A Machine Learning Algorithm for Finite-Horizon Stochastic Control Problems in Economics](https://www.ml-quant.com/papers/arxiv/2411.08668/): A proposed machine learning algorithm effectively solves complex, time-limited stochastic control problems, showing good convergence and efficiency without depending on the Bellman equation.
- [Neural Network Convergence for Variational Inequalities](https://www.ml-quant.com/papers/arxiv/2509.26535/): A novel approach to using neural networks on linear parabolic variational inequalities shows the potential of neural networks in solving optimal stopping and control problems in finance.
