<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://vaquero.me/atom.xml</id>
  <title>Lorenzo Vaquero — Publications</title>
  <subtitle>New papers from Lorenzo Vaquero on computer vision, multi-object tracking, human–object interaction, vision–language models, and AI explainability.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://vaquero.me/atom.xml"/>
  <link rel="alternate" type="text/html" href="https://vaquero.me/publications/"/>
  <updated>2026-01-01T00:00:00.000Z</updated>
  <author>
    <name>Lorenzo Vaquero Otal</name>
    <uri>https://vaquero.me/</uri>
  </author>
  <entry>
    <id>https://vaquero.me/publications/2026-cvpr-from-weights-to-concepts/</id>
    <title>From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-cvpr-from-weights-to-concepts/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>A data-free, training-free framework that directly analyses CLIP&apos;s vision transformer in weight space, decomposing each attention head into singular vectors linked to textual concepts. (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026)</summary>
    <author><name>Francesco Gentile</name></author>
    <author><name>Nicola Dall&apos;Asen</name></author>
    <author><name>Francesco Tonini</name></author>
    <author><name>Massimiliano Mancini</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Vision-Language Models"/>
    <category term="AI Explainability"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-perconai-efficient-portrait-segmentation-embedded-devices/</id>
    <title>Enabling Efficient Portrait Segmentation on Embedded Devices</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-perconai-efficient-portrait-segmentation-embedded-devices/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>SINergy is a scalable portrait segmentation network designed around embedded hardware constraints, reaching real-time segmentation on microcontroller-scale devices. (IEEE International Conference on Pervasive Computing and Communications Workshops, 2026)</summary>
    <author><name>Riccardo Benevelli</name></author>
    <author><name>Alberto Ancilotto</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Elisa Ricci</name></author>
    <author><name>Elisabetta Farella</name></author>
    <category term="Embedded Vision"/>
    <category term="Efficient Segmentation"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-acmmm-finehoi/</id>
    <title>FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-acmmm-finehoi/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>FineHOI detects unseen human-object interactions by reasoning over dense, part-level visual cues, improving fine-grained recognition without LLM or detector features. (ACM International Conference on Multimedia, 2026)</summary>
    <author><name>Francesco Tonini</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Mohammad Mahdi Derakhshani</name></author>
    <author><name>Cees G. M. Snoek</name></author>
    <author><name>Elisa Ricci</name></author>
    <author><name>Cigdem Beyan</name></author>
    <category term="Human-Object Interaction"/>
    <category term="Vision-Language Models"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-nature-humanitys-last-exam/</id>
    <title>Humanity&apos;s Last Exam: A Benchmark of Expert-Level Academic Questions to Assess AI Capabilities</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-nature-humanitys-last-exam/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>Humanity&apos;s Last Exam is a multimodal benchmark of 2,500 expert-level closed-ended academic questions designed to measure frontier AI capabilities beyond saturated benchmarks. (Nature, 2026)</summary>
    <author><name>Center for AI Safety</name></author>
    <author><name>Scale AI</name></author>
    <author><name>HLE Contributors Consortium</name></author>
    <category term="Benchmarks"/>
    <category term="AI Evaluation"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-caepia-keynote-lost-and-found/</id>
    <title>Keynote on Overcoming Detector Failures in Online Multi-Object Tracking</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-caepia-keynote-lost-and-found/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>This keynote paper summarizes BUSCA, a work published in ECCV&apos;24. BUSCA plugs into online tracking-by-detection systems to recover objects missed by detectors through proposal generation and decision-transformer association. (Conference of the Spanish Association for Artificial Intelligence, 2026)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Yihong Xu</name></author>
    <author><name>Xavier Alameda-Pineda</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Online Tracking"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-caepia-keynote-superpowering-open-vocabulary-object-detectors-x-ray-vision/</id>
    <title>Keynote on Superpowering Open-Vocabulary Object Detectors for X-ray Vision</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-caepia-keynote-superpowering-open-vocabulary-object-detectors-x-ray-vision/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>This keynote paper summarizes RAXO, a work published in ICCV&apos;25. RAXO adapts RGB open-vocabulary detectors to X-ray imagery with training-free visual descriptors and introduces DET-COMPASS for large-scale X-ray OvOD evaluation. (Conference of the Spanish Association for Artificial Intelligence, 2026)</summary>
    <author><name>Pablo Garcia-Fernandez</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Mingxuan Liu</name></author>
    <author><name>Feng Xue</name></author>
    <author><name>Daniel Cores</name></author>
    <author><name>Nicu Sebe</name></author>
    <author><name>Manuel Mucientes</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Open-Vocabulary Detection"/>
    <category term="X-ray Vision"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-fg-towards-unconstrained-human-object-interaction/</id>
    <title>Towards Unconstrained Human-Object Interaction</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-fg-towards-unconstrained-human-object-interaction/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>AnyHOI studies human-object interaction detection without a predefined interaction vocabulary, using multimodal LLMs and language-to-graph conversion to recover structured triplets from free-form outputs. (IEEE International Conference on Automatic Face and Gesture Recognition, 2026)</summary>
    <author><name>Francesco Tonini</name></author>
    <author><name>Alessandro Conti</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Cigdem Beyan</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Human-Object Interaction"/>
    <category term="Multimodal LLMs"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-fg-training-free-semantic-multi-object-tracking/</id>
    <title>Training-Free Semantic Multi-Object Tracking with Vision-Language Models</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-fg-training-free-semantic-multi-object-tracking/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>TF-SMOT composes frozen detection, segmentation tracking, video-language generation, and LLM disambiguation modules to produce semantic tracking outputs without task-specific training. (IEEE International Conference on Automatic Face and Gesture Recognition, 2026)</summary>
    <author><name>Laurence Bonat</name></author>
    <author><name>Francesco Tonini</name></author>
    <author><name>Elisa Ricci</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Vision-Language Models"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2026-fg-zero-shot-temporal-action-localization/</id>
    <title>Zero-Shot Temporal Action Localization Through Textual Guidance</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2026-fg-zero-shot-temporal-action-localization/"/>
    <updated>2026-01-01T00:00:00.000Z</updated>
    <published>2026-01-01T00:00:00.000Z</published>
    <summary>TeGu compensates for the lack of supervision in zero-shot temporal action localization by exploiting rich textual cues from large language models, distinguishing foreground from background frames without training. (IEEE International Conference on Automatic Face and Gesture Recognition, 2026)</summary>
    <author><name>Benedetta Liberatori</name></author>
    <author><name>Alessandro Conti</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Paolo Rota</name></author>
    <author><name>Yiming Wang</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Temporal Action Localization"/>
    <category term="Vision-Language Models"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2025-neurips-convis-bench/</id>
    <title>ConViS-Bench: Estimating Video Similarity Through Semantic Concepts</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2025-neurips-convis-bench/"/>
    <updated>2025-01-01T00:00:00.000Z</updated>
    <published>2025-01-01T00:00:00.000Z</published>
    <summary>ConViS-Bench evaluates video similarity along semantic concepts, pairing concept-level human scores with free-form descriptions for interpretable video comparison. (Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025)</summary>
    <author><name>Benedetta Liberatori</name></author>
    <author><name>Alessandro Conti</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Yiming Wang</name></author>
    <author><name>Elisa Ricci</name></author>
    <author><name>Paolo Rota</name></author>
    <category term="Video Understanding"/>
    <category term="Benchmarks"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2025-acmmm-dynamic-scoring-hoi-detection/</id>
    <title>Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2025-acmmm-dynamic-scoring-hoi-detection/"/>
    <updated>2025-01-01T00:00:00.000Z</updated>
    <published>2025-01-01T00:00:00.000Z</published>
    <summary>DYSCO is a training-free HOI detector that combines textual and visual interaction representations in a multimodal registry with dynamic multi-head scoring. (ACM International Conference on Multimedia, 2025)</summary>
    <author><name>Francesco Tonini</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Alessandro Conti</name></author>
    <author><name>Cigdem Beyan</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Human-Object Interaction"/>
    <category term="Vision-Language Models"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2025-ibpria-enhancing-multi-object-tracking-segmentation-masks/</id>
    <title>Enhancing Multi-Object Tracking with Segmentation Masks: A Solution for Lost Object Recovery</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2025-ibpria-enhancing-multi-object-tracking-segmentation-masks/"/>
    <updated>2025-01-01T00:00:00.000Z</updated>
    <published>2025-01-01T00:00:00.000Z</published>
    <summary>A lost-track recovery architecture uses segmentation masks and a transformer-based mask selector to preserve tracks when detectors fail in crowded scenes. (Iberian Conference on Pattern Recognition and Image Analysis, 2025)</summary>
    <author><name>Manuel Bendaña</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Segmentation"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2025-iccv-superpowering-open-vocabulary-object-detectors-x-ray-vision/</id>
    <title>Superpowering Open-Vocabulary Object Detectors for X-ray Vision</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2025-iccv-superpowering-open-vocabulary-object-detectors-x-ray-vision/"/>
    <updated>2025-01-01T00:00:00.000Z</updated>
    <published>2025-01-01T00:00:00.000Z</published>
    <summary>RAXO adapts RGB open-vocabulary detectors to X-ray imagery with training-free visual descriptors and introduces DET-COMPASS for large-scale X-ray OvOD evaluation. (IEEE/CVF International Conference on Computer Vision, 2025)</summary>
    <author><name>Pablo Garcia-Fernandez</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Mingxuan Liu</name></author>
    <author><name>Feng Xue</name></author>
    <author><name>Daniel Cores</name></author>
    <author><name>Nicu Sebe</name></author>
    <author><name>Manuel Mucientes</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Open-Vocabulary Detection"/>
    <category term="X-ray Vision"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2024-eccv-lost-and-found/</id>
    <title>Lost and Found: Overcoming Detector Failures in Online Multi-Object Tracking</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2024-eccv-lost-and-found/"/>
    <updated>2024-01-01T00:00:00.000Z</updated>
    <published>2024-01-01T00:00:00.000Z</published>
    <summary>BUSCA plugs into online tracking-by-detection systems to recover objects missed by detectors through proposal generation and decision-transformer association. (European Conference on Computer Vision, 2024)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Yihong Xu</name></author>
    <author><name>Xavier Alameda-Pineda</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Online Tracking"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2024-acmmm-al-gtd/</id>
    <title>AL-GTD: Deep Active Learning for Gaze Target Detection</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2024-acmmm-al-gtd/"/>
    <updated>2024-01-01T00:00:00.000Z</updated>
    <published>2024-01-01T00:00:00.000Z</published>
    <summary>AL-GTD reduces gaze target annotation cost by combining supervised and self-supervised signals in an active learning acquisition function with pseudo-labeling. (ACM International Conference on Multimedia, 2024)</summary>
    <author><name>Francesco Tonini</name></author>
    <author><name>Nicola Dall&apos;Asen</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Cigdem Beyan</name></author>
    <author><name>Elisa Ricci</name></author>
    <category term="Gaze Target Detection"/>
    <category term="Active Learning"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2024-iscas-live-demonstration-sram-dnn-cim/</id>
    <title>Live Demonstration: 5-bit signed SRAM-based DNN CIM for Image Recognition</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2024-iscas-live-demonstration-sram-dnn-cim/"/>
    <updated>2024-01-01T00:00:00.000Z</updated>
    <published>2024-01-01T00:00:00.000Z</published>
    <summary>A live mixed-signal compute-in-memory demonstration uses 5-bit signed SRAM weights and PWM-coded images for low-power DNN image recognition. (IEEE International Symposium on Circuits and Systems, 2024)</summary>
    <author><name>Ó. Pereira-Rial</name></author>
    <author><name>D. García-Lesta</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>P. López</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>D. Cabello</name></author>
    <category term="Compute-in-Memory"/>
    <category term="Embedded AI"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2023-pattern-recognition-real-time-siamese-multiple-object-tracker/</id>
    <title>Real-Time Siamese Multiple Object Tracker with Enhanced Proposals</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2023-pattern-recognition-real-time-siamese-multiple-object-tracker/"/>
    <updated>2023-01-01T00:00:00.000Z</updated>
    <published>2023-01-01T00:00:00.000Z</published>
    <summary>SiamMOTION tracks dozens of arbitrary objects in real time through feature-pyramid proposals, attention, inertia, pairwise depthwise RPN matching, and multi-object penalization. (Pattern Recognition, 2023)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Real-Time Vision"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2023-tci-depth-estimation-image-restoration-defocused-images/</id>
    <title>Depth Estimation and Image Restoration by Deep Learning from Defocused Images</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2023-tci-depth-estimation-image-restoration-defocused-images/"/>
    <updated>2023-01-01T00:00:00.000Z</updated>
    <published>2023-01-01T00:00:00.000Z</published>
    <summary>2HDED:NET jointly estimates depth and restores all-in-focus images from defocused input through a shared encoder and parallel task branches. (IEEE Transactions on Computational Imaging, 2023)</summary>
    <author><name>Saqib Nazir</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Manuel Mucientes</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Daniela Coltuc</name></author>
    <category term="Depth Estimation"/>
    <category term="Image Restoration"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2022-icip-2hded-net/</id>
    <title>2HDED:NET for Joint Depth Estimation and Image Deblurring from a Single Out-of-Focus Image</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2022-icip-2hded-net/"/>
    <updated>2022-01-01T00:00:00.000Z</updated>
    <published>2022-01-01T00:00:00.000Z</published>
    <summary>2HDED:NET performs depth-from-defocus and all-in-focus image restoration in parallel through a shared encoder and two balanced decoder branches. (IEEE International Conference on Image Processing, 2022)</summary>
    <author><name>Saqib Nazir</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Manuel Mucientes</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Daniela Coltuc</name></author>
    <category term="Depth Estimation"/>
    <category term="Image Deblurring"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2022-icpr-fast-multi-object-tracking-feature-pyramid-region-proposal-networks/</id>
    <title>Fast Multi-Object Tracking with Feature Pyramid and Region Proposal Networks</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2022-icpr-fast-multi-object-tracking-feature-pyramid-region-proposal-networks/"/>
    <updated>2022-01-01T00:00:00.000Z</updated>
    <published>2022-01-01T00:00:00.000Z</published>
    <summary>SiamFAST brings feature-pyramid ROI extraction, pairwise depthwise RPN matching, and multi-object penalization to real-time arbitrary multi-object tracking. (International Conference on Pattern Recognition, 2022)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Real-Time Vision"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2021-pattern-recognition-tracking-more-than-100-arbitrary-objects/</id>
    <title>Tracking More Than 100 Arbitrary Objects at 25 FPS Through Deep Learning</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2021-pattern-recognition-tracking-more-than-100-arbitrary-objects/"/>
    <updated>2021-01-01T00:00:00.000Z</updated>
    <published>2021-01-01T00:00:00.000Z</published>
    <summary>SiamMT scales visual object tracking to dozens of arbitrary objects by reusing feature computations and introducing efficient crop-resize and pairwise similarity operators. (Pattern Recognition, 2021)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Deep Learning"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2020-icpr-siammt/</id>
    <title>SiamMT: Real-Time Arbitrary Multi-Object Tracking</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2020-icpr-siammt/"/>
    <updated>2020-01-01T00:00:00.000Z</updated>
    <published>2020-01-01T00:00:00.000Z</published>
    <summary>SiamMT is a Siamese convolutional architecture for tracking multiple arbitrary objects in real time by sharing frame features and using pairwise cross-correlation. (International Conference on Pattern Recognition, 2020)</summary>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Manuel Mucientes</name></author>
    <author><name>Víctor M. Brea</name></author>
    <category term="Multi-Object Tracking"/>
    <category term="Siamese Networks"/>
  </entry>
  <entry>
    <id>https://vaquero.me/publications/2018-wgml-deep-learning-video-object-detection-tracking/</id>
    <title>Deep Learning for Video Object Detection and Tracking</title>
    <link rel="alternate" type="text/html" href="https://vaquero.me/publications/2018-wgml-deep-learning-video-object-detection-tracking/"/>
    <updated>2018-01-01T00:00:00.000Z</updated>
    <published>2018-01-01T00:00:00.000Z</published>
    <summary>Brief summary of deep convolutional approaches for small-object detection, real-time multi-object tracking, and integrated traffic monitoring. (Machine Learning Workshop Galicia, 2018)</summary>
    <author><name>Brais Bosquet</name></author>
    <author><name>Mauro Fernández-Sanjurjo</name></author>
    <author><name>Lorenzo Vaquero</name></author>
    <author><name>Daniel Cores</name></author>
    <author><name>Víctor M. Brea</name></author>
    <author><name>Manuel Mucientes</name></author>
    <category term="Object Detection"/>
    <category term="Multi-Object Tracking"/>
  </entry>
</feed>
