Compositional Video Understanding with Spatiotemporal Structure-based Transformers

Hoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol Kim
2/9/2026
DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. Any repos shown below are weak matches — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{yun2026compositional,
  title  = {Compositional Video Understanding with Spatiotemporal Structure-based Transformers},
  author = {Hoyeoung Yun and Jinwoo Ahn and Minseo Kim and Eun-Sol Kim},
  year   = {2026},
  doi    = {10.1109/CVPR52733.2024.01774},
  url    = {https://doi.org/10.1109/CVPR52733.2024.01774},
  journal = {CVPR 2024 2024}
}

Discussion