---
title: MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
url: https://www.ml-quant.com/papers/arxiv/2410.02743/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2410.02743
source_url: https://arxiv.org/abs/2410.02743
featured: 2024-10-09
citations: 15
topic: ML & AI Methods
---


# MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

The MA-RLHF framework integrates macro actions into the learning process of large language models, enhancing learning efficiency and performance in tasks like text summarization and dialogue generation.

- Source: https://arxiv.org/abs/2410.02743
- Identifier: arXiv:2410.02743
- Released: 2024-10-03
- First featured: Quant Letter No. 69 (2024-10-09): https://www.ml-quant.com/issues/2024-10-09/
- Citations (Semantic Scholar): 15
- Published in: International Conference on Learning Representations
- Topic: ML & AI Methods

## Related

- [VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback](https://www.ml-quant.com/papers/arxiv/2409.18417/): An auction mechanism is introduced to enhance cost-efficiency in fine-tuning large language models using Reinforcement Learning from Human Feedback, focusing on quality feedback and model performance.
- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
- [Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia](https://www.ml-quant.com/papers/arxiv/2312.03664/): Concordia is a library designed to help build and operate Generative Agent-Based Models (GABMs), using Large Language Models (LLMs) to simulate physical or digital environments.
