Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection
Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci
ACM International Conference on Multimedia
DYSCO is a training-free HOI detector that combines textual and visual interaction representations in a multimodal registry with dynamic multi-head scoring.
Human-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues, but these annotations are labor-intensive, inconsistent, and difficult to scale to new domains and rare interactions. We propose Dynamic Scoring with enhanced semantics (DYSCO), a training-free HOI detection framework that uses textual and visual interaction representations within a multimodal registry for robust interaction understanding. The registry incorporates a small set of visual cues and interaction signatures to improve semantic alignment of verbs, and a multi-head attention mechanism adaptively weights visual and textual features for each test sample. DYSCO surpasses training-free state-of-the-art models and is competitive with training-based approaches, particularly on rare interactions.
@inproceedings{tonini2025dynamic,
author = {Francesco Tonini and
Lorenzo Vaquero and
Alessandro Conti and
Cigdem Beyan and
Elisa Ricci},
title = {Dynamic Scoring with Enhanced Semantics for Training-Free
Human-Object Interaction Detection},
booktitle = {Proc. {ACM} Int. Conf. Multimedia ({ACM MM})},
pages = {2801-2810},
year = {2025},
doi = {10.1145/3746027.3754770}
}
