ML-QuantSubscribe

Machine learningLLMs & Text

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Vision and Speech Interaction: A proposed training methodology enables Large Language Models to comprehend visual and speech data, improving speech-to-speech dialogue capabilities and response speed.

Featured in No. 81 on 8 Jan 2025 · 5 days after release · 230 citations today · published in Neural Information Processing Systems

Released
3 Jan 2025
First featured
No. 81 · 8 Jan 2025
Citations (Semantic Scholar)
230
Influential citations
26
Published in
Neural Information Processing Systems
Shares when featured
52
Identifier
arXiv:2501.01957

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page