Machine learningML & AI Methods
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
The study reveals that scaling the decoder in auto-encoders, specifically the VisionTransformer architecture for Tokenization (ViTok), improves reconstruction performance and sets new standards for class-conditional video generation when combined with Diffusion Transformers.
Featured in No. 83 on 23 Jan 2025 · 7 days after release · 33 citations today · published in International Conference on Machine Learning
- Released
- 16 Jan 2025
- First featured
- No. 83 · 23 Jan 2025
- Citations (Semantic Scholar)
- 33
- Influential citations
- 1
- Published in
- International Conference on Machine Learning
- Shares when featured
- 31
- Identifier
- arXiv:2501.09755
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).