---
title: HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation
url: https://www.ml-quant.com/papers/arxiv/2407.17438/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2407.17438
source_url: https://arxiv.org/pdf/2407.17438
featured: 2024-07-31
citations: 74
topic: Other
---


# HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation

Camera-controllable Human Image Animation: HumanVid, a new large-scale dataset for human image animation that combines real and synthetic data, has been developed by researchers, setting a new standard in the field.

- Source: https://arxiv.org/pdf/2407.17438
- Identifier: arXiv:2407.17438
- Released: 2024-07-24
- First featured: Quant Letter No. 59 (2024-07-31): https://www.ml-quant.com/issues/2024-07-31/
- Citations (Semantic Scholar): 74
- Published in: Neural Information Processing Systems
- Topic: Other

## Related

- [Scaling Laws of Synthetic Images for Model Training … for Now](https://www.ml-quant.com/papers/arxiv/2312.04567/): The study investigates the scaling laws of synthetic images used in training supervised models, identifying factors that influence scaling behavior and situations where scaling synthetic data is most effective.
- [R.I.P.: Better Models by Survival of the Fittest Prompts](https://www.ml-quant.com/papers/arxiv/2501.18578/): The study introduces Rejecting Instruction Preferences (RIP), a method for evaluating data integrity that can filter prompts or create synthetic datasets, enhancing performance across various benchmarks.
- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [Qwen2.5-Coder Technical Report](https://www.ml-quant.com/papers/arxiv/2409.12186/): The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
- [Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives](https://www.ml-quant.com/papers/arxiv/2311.18259/): Understanding Human Activity: The paper presents Ego-Exo4D, a large-scale video dataset and benchmark challenge featuring human activities from various perspectives, aimed at improving first-person video understanding.
