AL-GTD: Deep Active Learning for Gaze Target Detection
Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci
ACM International Conference on Multimedia
AL-GTD reduces gaze target annotation cost by combining supervised and self-supervised signals in an active learning acquisition function with pseudo-labeling.
Gaze target detection aims to determine the image location where a person is looking. Existing studies have made significant progress by regressing accurate gaze heatmaps, but these results largely rely on extensive labeled datasets that demand substantial human labor. AL-GTD reduces this reliance by integrating supervised and self-supervised losses within a sample acquisition function for active learning. It also uses pseudo-labeling to mitigate distribution shifts during training. AL-GTD achieves the best AUC results by using only 40-50% of the training data, while state-of-the-art gaze target detectors require the entire training set to reach the same performance. It also reaches satisfactory performance with 10-20% of the training data, showing that the acquisition function selects informative samples effectively.
@inproceedings{tonini2024algtd,
author = {Francesco Tonini and
Nicola Dall'Asen and
Lorenzo Vaquero and
Cigdem Beyan and
Elisa Ricci},
title = {{AL-GTD}: Deep Active Learning for Gaze Target Detection},
booktitle = {Proc. {ACM} Int. Conf. Multimedia ({ACM MM})},
pages = {2360-2369},
year = {2024},
doi = {10.1145/3664647.3680952}
}
