Skip to content
Cover of Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci

ACM International Conference on Multimedia

DYSCO is a training-free HOI detector that combines textual and visual interaction representations in a multimodal registry with dynamic multi-head scoring.

PDFCode

Human-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues, but these annotations are labor-intensive, inconsistent, and difficult to scale to new domains and rare interactions. We propose Dynamic Scoring with enhanced semantics (DYSCO), a training-free HOI detection framework that uses textual and visual interaction representations within a multimodal registry for robust interaction understanding. The registry incorporates a small set of visual cues and interaction signatures to improve semantic alignment of verbs, and a multi-head attention mechanism adaptively weights visual and textual features for each test sample. DYSCO surpasses training-free state-of-the-art models and is competitive with training-based approaches, particularly on rare interactions.

@inproceedings{tonini2025dynamic,
  author    = {Francesco Tonini and
               Lorenzo Vaquero and
               Alessandro Conti and
               Cigdem Beyan and
               Elisa Ricci},
  title     = {Dynamic Scoring with Enhanced Semantics for Training-Free
               Human-Object Interaction Detection},
  booktitle = {Proc. {ACM} Int. Conf. Multimedia ({ACM MM})},
  pages     = {2801-2810},
  year      = {2025},
  doi       = {10.1145/3746027.3754770}
}

Click the image to zoom · drag to pan · ESC to close