Machine learningLLMs & Text
BLINK: Multimodal Large Language Models Can See but Not Perceive
Multimodal LLMs Visual Perception Benchmark: The authors present Blink, a benchmark for multimodal language models that tests visual perception abilities, showing that current models struggle with these tasks.
Featured in No. 46 on 24 Apr 2024 · 6 days after release · 603 citations today · published in European Conference on Computer Vision
- Released
- 18 Apr 2024
- First featured
- No. 46 · 24 Apr 2024
- Citations (Semantic Scholar)
- 603
- Influential citations
- 78
- Published in
- European Conference on Computer Vision
- Shares when featured
- 21
- Identifier
- arXiv:2404.12390
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).