---
title: Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
url: https://www.ml-quant.com/papers/arxiv/2408.00754/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2408.00754
source_url: https://arxiv.org/pdf/2408.00754.pdf
featured: 2024-08-07
citations: 20
topic: LLMs & Text
---


# Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

The paper presents Coarse Correspondence, a visual prompting method that enhances multimodal language models' understanding of 3D and temporal dimensions, achieving top results on various benchmarks.

- Source: https://arxiv.org/pdf/2408.00754.pdf
- Identifier: arXiv:2408.00754
- Released: 2024-08-01
- First featured: Quant Letter No. 60 (2024-08-07): https://www.ml-quant.com/issues/2024-08-07/
- Citations (Semantic Scholar): 20
- Published in: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- Topic: LLMs & Text

## Related

- [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://www.ml-quant.com/papers/arxiv/2402.03300/): Advancing Math Reasoning in Language Models: DeepSeekMath7B is a new language model that uses web data and Group Relative Policy Optimization for advanced mathematical reasoning, scoring high on the MATH benchmark.
- [Mistral 7B](https://www.ml-quant.com/papers/arxiv/2310.06825/): Superior Language Model: Mistral 7B v0.1 is a language model with 7 billion parameters that excels in reasoning, mathematics, and code generation, and has a version specifically designed to follow instructions.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model](https://www.ml-quant.com/papers/arxiv/2405.04434/): MoE Language Model: DeepSeek-V2, a language model with 236B parameters, offers enhanced performance and cost efficiency compared to its predecessor, ranking high among open-source models.
