---
title: SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
url: https://www.ml-quant.com/papers/arxiv/2501.18564/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2501.18564
source_url: https://arxiv.org/abs/2501.18564
featured: 2025-02-05
citations: 89
topic: Other
---


# SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation

Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.

- Source: https://arxiv.org/abs/2501.18564
- Identifier: arXiv:2501.18564
- Released: 2025-01-30
- First featured: Quant Letter No. 84 (2025-02-05): https://www.ml-quant.com/issues/2025-02-05/
- Citations (Semantic Scholar): 89
- Published in: International Conference on Machine Learning
- Topic: Other

## Related

- [Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction](https://www.ml-quant.com/papers/arxiv/2409.18124/): Visual Foundation for Dense Prediction: Lotus, a new visual foundation model, predicts annotations directly, improving inference speed and performance in zero-shot depth and normal estimation tasks.
- [A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities](https://www.ml-quant.com/papers/arxiv/2405.12971/): Image Parsing Model: BiomedParse is a new tool for biomedical image analysis, capable of identifying 82 object types across 9 imaging modalities, enhancing accuracy in biomedical research.
- [OmniGlue: Generalizable Feature Matching with Foundation Model Guidance](https://www.ml-quant.com/papers/arxiv/2405.12979/): OmniGlue, a new image matcher that performs better on unseen image domains than previous models, is introduced in this paper.
- [Foundation Models for Music: A Survey](https://www.ml-quant.com/papers/arxiv/2408.14340/): The article discusses the influence of foundation models on the music industry, emphasizing their potential in music generation and the need for ethical research on issues like transparency and copyright.
- [Feat2GS: Probing Visual Foundation Models with Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2412.09606/): Visual Foundation Models: Feat2GS is a framework that extracts 3D Gaussians attributes from unposed images, allowing for the probing of 3D awareness for geometry and texture through novel view synthesis.
- [LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation](https://www.ml-quant.com/papers/arxiv/2501.16559/): Model Adaptation: LoRA-X enables the transfer of fine-tuning parameters across different models, enhancing the efficiency of text-to-image generation without needing original training data.
