---
title: InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma
url: https://www.ml-quant.com/papers/arxiv/2411.09856/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2411.09856
source_url: http://arxiv.org/abs/2411.09856v2
featured: 2024-12-04
citations: 3
topic: ML & AI Methods
---


# InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma

InvestESG is a new benchmark using advanced learning to study the effects of ESG disclosure mandates on corporate climate investments, indicating that ESG-aware investors can boost corporate cooperation and mitigate climate risks.

- Source: http://arxiv.org/abs/2411.09856v2
- Identifier: arXiv:2411.09856
- Released: 2024-11-15
- First featured: Quant Letter No. 77 (2024-12-04): https://www.ml-quant.com/issues/2024-12-04/
- Citations (Semantic Scholar): 3
- Published in: International Conference on Learning Representations
- Topic: ML & AI Methods

## Related

- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [DPO Meets PPO: Reinforced Token Optimization for RLHF](https://www.ml-quant.com/papers/arxiv/2404.18922/): A new framework is introduced that models Reinforcement Learning from Human Feedback as a Markov decision process, using an algorithm that learns from preference data.
- [Settling the Sample Complexity of Model-Based Offline Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2204.05275/): A paper reveals that a model-based approach can achieve optimal sample complexity without burn-in cost in offline reinforcement learning for tabular Markov decision processes, providing an efficient solution for sample-starved applications.
- [Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data](https://www.ml-quant.com/papers/arxiv/2412.07762/): The article introduces Warm-start RL (WSRL), a new reinforcement learning approach that doesn't require offline data, leading to quicker learning and better performance than previous algorithms.
- [Empirical Design in Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2304.01315/): The article highlights the importance of proper statistical evidence and avoiding common errors in empirical design for effective reinforcement learning experiments.
