Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data

Abstract

A new evaluation task for video-text understanding requiring cross-frame reasoning, and an LLM-teacher approach to learn discriminative action embeddings.

Publication
ECCV 2024: 254-269