ML-QuantSubscribe

Machine learningML & AI Methods

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

The study reveals that scaling the decoder in auto-encoders, specifically the VisionTransformer architecture for Tokenization (ViTok), improves reconstruction performance and sets new standards for class-conditional video generation when combined with Diffusion Transformers.

Featured in No. 83 on 23 Jan 2025 · 7 days after release · 33 citations today · published in International Conference on Machine Learning

Released
16 Jan 2025
First featured
No. 83 · 23 Jan 2025
Citations (Semantic Scholar)
33
Influential citations
1
Published in
International Conference on Machine Learning
Shares when featured
31
Identifier
arXiv:2501.09755

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page