---
title: Moving Object Segmentation: All You Need Is SAM (and Flow)
url: https://www.ml-quant.com/papers/arxiv/2404.12389/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2404.12389
source_url: https://arxiv.org/pdf/2404.12389.pdf
featured: 2024-04-24
citations: 39
topic: Other
---


# Moving Object Segmentation: All You Need Is SAM (and Flow)

The study applies the Segment Anything model (SAM) to motion segmentation in videos, showing that simple methods combining SAM with optical flow surpass previous approaches.

- Source: https://arxiv.org/pdf/2404.12389.pdf
- Identifier: arXiv:2404.12389
- Released: 2024-04-18
- First featured: Quant Letter No. 46 (2024-04-24): https://www.ml-quant.com/issues/2024-04-24/
- Citations (Semantic Scholar): 39
- Published in: Asian Conference on Computer Vision
- Topic: Other

## Related

- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [Qwen2.5-Coder Technical Report](https://www.ml-quant.com/papers/arxiv/2409.12186/): The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
- [Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives](https://www.ml-quant.com/papers/arxiv/2311.18259/): Understanding Human Activity: The paper presents Ego-Exo4D, a large-scale video dataset and benchmark challenge featuring human activities from various perspectives, aimed at improving first-person video understanding.
- [Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2310.10642/): The 4DGS model is introduced, capable of reconstructing dynamic 3D scenes from 2D images and generating diverse views over time, providing real-time rendering efficiency.
- [FateZero: Fusing Attentions for Zero-shot Text-based Video Editing](https://www.ml-quant.com/papers/arxiv/2303.09535/): Text-based Video Editing: The article introduces FateZero, a new method for editing real-world videos using text, which outperforms previous models in maintaining video structure, motion, and frame consistency.
