---
title: How Far is Video Generation from World Model: A Physical Law Perspective
url: https://www.ml-quant.com/papers/arxiv/2411.02385/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2411.02385
source_url: https://arxiv.org/abs/2411.02385
featured: 2024-11-06
citations: 230
topic: Other
---


# How Far is Video Generation from World Model: A Physical Law Perspective

A Physical Law Perspective: The paper assesses the capability of video generation models to identify physical laws from visual data, indicating that mere scaling is not enough for these models to discover fundamental physical laws.

- Source: https://arxiv.org/abs/2411.02385
- Identifier: arXiv:2411.02385
- Released: 2024-11-04
- First featured: Quant Letter No. 73 (2024-11-06): https://www.ml-quant.com/issues/2024-11-06/
- Citations (Semantic Scholar): 230
- Published in: International Conference on Machine Learning
- Topic: Other

## Related

- [Depth Anything V2](https://www.ml-quant.com/papers/arxiv/2406.09414/): Depth Anything V2 is a new model for monocular depth estimation, using synthetic and large-scale pseudo-labeled real images for faster, more accurate results and setting a new evaluation benchmark.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [Qwen2.5-Coder Technical Report](https://www.ml-quant.com/papers/arxiv/2409.12186/): The report unveils the Qwen2.5-Coder series, an improvement from its predecessor, showcasing remarkable code generation abilities and achieving top-tier performance in various code-related tasks.
- [Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives](https://www.ml-quant.com/papers/arxiv/2311.18259/): Understanding Human Activity: The paper presents Ego-Exo4D, a large-scale video dataset and benchmark challenge featuring human activities from various perspectives, aimed at improving first-person video understanding.
- [Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2310.10642/): The 4DGS model is introduced, capable of reconstructing dynamic 3D scenes from 2D images and generating diverse views over time, providing real-time rendering efficiency.
- [Continuous 3D Perception Model with Persistent State](https://www.ml-quant.com/papers/arxiv/2501.12387/): The paper presents CUT3R, a unified framework that uses a recurrent model to generate metric-scale pointmaps from a stream of images, enabling dense scene reconstruction that updates with new images.
