Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language

Chuanhao Li, Zhen Li, Chenchen Jing, Yunde Jia, Yuwei Wu
2/14/2026

Abstract

Compositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand the effect of the primitives, including words, image regions, and video frames, to improve the compositional generalization capability. In this paper, we explore the effect of primitives for compositional generalization in V&L. Specifically, we present a self-supervised learning based framework that equips existing V&L methods with two characteristics: semantic equivariance and semantic invariance. With the two characteristics, the methods understand primitives by perceiving the effect of primitive changes on sample semantics and ground-truth. Experimental results on two tasks: temporal video grounding and visual question answering, demonstrate the effectiveness of our framework.

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 1 weaker match is hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{li2026exploring,
  title  = {Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language},
  author = {Chuanhao Li and Zhen Li and Chenchen Jing and Yunde Jia and Yuwei Wu},
  year   = {2026},
  doi    = {10.1109/CVPR52729.2023.01830},
  url    = {https://doi.org/10.1109/CVPR52729.2023.01830},
  journal = {CVPR 2023 2023}
}

Discussion