Machine learningLLMs & Text
Wolf: Dense Video Captioning with a World Summarization Framework
Wolf, a new video captioning framework, uses Vision Language Models to efficiently summarize information, outperforming existing methods and setting a new standard for video captioning.
Featured in No. 59 on 31 Jul 2024 · 5 days after release · 6 citations today · published in Trans. Mach. Learn. Res.
- Released
- 26 Jul 2024
- First featured
- No. 59 · 31 Jul 2024
- Citations (Semantic Scholar)
- 6
- Influential citations
- 0
- Published in
- Trans. Mach. Learn. Res.
- Shares when featured
- 9
- Identifier
- arXiv:2407.18908
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).