When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and Approach

Tao Ma, Bing Bai, Haozhe Lin, Heyuan Wang, Yu Wang +2 more
2/9/2026

Abstract

Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless, recent advances in imaging technology have enabled the acquisition of gigapixel-level images, providing high-resolution details in large-scale scenes containing numerous objects. To bridge this gap between imaging and computer vision benchmarks and make grounding more practically valuable, we introduce a novel dataset, named GigaGrounding, designed to challenge visual grounding models in gigapixel-level large-scale scenes. We extensively analyze and compare the dataset with existing benchmarks, demonstrating that GigaGrounding presents unique challenges such as large-scale scene understanding, gigapixel-level resolution, significant variations in object scales, and the “multi-hop expressions”. Furthermore, we introduced a simple yet effective grounding approach, which employs a “glance-to-zoom-in” paradigm and exhibits enhanced capabilities for addressing the GigaGrounding task. The dataset is available at www.gigavision.ai.

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. Any repos shown below are weak matches — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{ma2026when,
  title  = {When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and Approach},
  author = {Tao Ma and Bing Bai and Haozhe Lin and Heyuan Wang and Yu Wang and Lin Luo and Lu Fang},
  year   = {2026},
  doi    = {10.1109/CVPR52733.2024.02088},
  url    = {https://doi.org/10.1109/CVPR52733.2024.02088},
  journal = {CVPR 2024 2024}
}

Discussion