ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang, Elisa Ricci, Paolo Rota
Conference on Neural Information Processing Systems Datasets and Benchmarks Track
ConViS-Bench evaluates video similarity along semantic concepts, pairing concept-level human scores with free-form descriptions for interpretable video comparison.
What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not been thoroughly studied and presents a challenge for models that often depend on broad global similarity scores. We introduce Concept-based Video Similarity estimation (ConViS), a task that compares pairs of videos by computing interpretable similarity scores across a predefined set of key semantic concepts. To support this task, we introduce ConViS-Bench, a benchmark comprising carefully annotated video pairs spanning multiple domains. Each pair comes with concept-level similarity scores and textual descriptions of both differences and similarities. We benchmark several state-of-the-art models on ConViS, providing insights into their alignment with human judgments and highlighting concept-specific challenges in video similarity estimation.
@inproceedings{liberatori2025convisbench,
author = {Benedetta Liberatori and
Alessandro Conti and
Lorenzo Vaquero and
Yiming Wang and
Elisa Ricci and
Paolo Rota},
title = {{ConViS-Bench}: Estimating Video Similarity Through Semantic
Concepts},
booktitle = {The Thirty-Ninth Annual Conference on Neural Information
Processing Systems Datasets and Benchmarks Track},
volume = {38},
year = {2025}
}
