Skip to content
Cover of ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

ConViS-Bench: Estimating Video Similarity Through Semantic Concepts

Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang, Elisa Ricci, Paolo Rota

Conference on Neural Information Processing Systems Datasets and Benchmarks Track

ConViS-Bench evaluates video similarity along semantic concepts, pairing concept-level human scores with free-form descriptions for interpretable video comparison.

What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not been thoroughly studied and presents a challenge for models that often depend on broad global similarity scores. We introduce Concept-based Video Similarity estimation (ConViS), a task that compares pairs of videos by computing interpretable similarity scores across a predefined set of key semantic concepts. To support this task, we introduce ConViS-Bench, a benchmark comprising carefully annotated video pairs spanning multiple domains. Each pair comes with concept-level similarity scores and textual descriptions of both differences and similarities. We benchmark several state-of-the-art models on ConViS, providing insights into their alignment with human judgments and highlighting concept-specific challenges in video similarity estimation.

@inproceedings{liberatori2025convisbench,
  author    = {Benedetta Liberatori and
               Alessandro Conti and
               Lorenzo Vaquero and
               Yiming Wang and
               Elisa Ricci and
               Paolo Rota},
  title     = {{ConViS-Bench}: Estimating Video Similarity Through Semantic
               Concepts},
  booktitle = {The Thirty-Ninth Annual Conference on Neural Information
               Processing Systems Datasets and Benchmarks Track},
  volume = {38},
  year      = {2025}
}

Click the image to zoom · drag to pan · ESC to close