TransVLAD: Multi-Scale Attention-Based Global Descriptors for Visual Geo-Localization

  • Yifan Xu*
  • , Pourya Shamsolmoali
  • , Eric Granger
  • , Claire Nicodeme
  • , Laurent Gardes
  • , Jie Yang
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

32 Scopus citations

Abstract

Visual geo-localization remains a challenging task due to variations in the appearance and perspective among captured images. This paper introduces an efficient TransVLAD module, which aggregates attention-based feature maps into a discriminative and compact global descriptor. Unlike existing methods that generate feature maps using only convolutional neural networks (CNNs), we propose a sparse transformer to encode global dependencies and compute attention-based feature maps, which effectively reduces visual ambiguities that occurs in large-scale geo-localization problems. A positional embedding mechanism is used to learn the corresponding geometric configurations between query and gallery images. A grouped VLAD layer is also introduced to reduce the number of parameters, and thus construct an efficient module. Finally, rather than only learning from the global descriptors on entire images, we propose a self-supervised learning method to further encode more information from multi-scale patches between the query and positive gallery images. Extensive experiments on three challenging large-scale datasets indicate that our model outperforms state-of-the-art models, and has lower computational complexity. The code is available at: https://github.com/wacv-23/TVLAD.

Original languageEnglish
Title of host publicationProceedings - 2023 IEEE Winter Conference on Applications of Computer Vision, WACV 2023
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2839-2848
Number of pages10
ISBN (Electronic)9781665493468
DOIs
StatePublished - 2023
Event23rd IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2023 - Waikoloa, United States
Duration: 3 Jan 20237 Jan 2023

Publication series

NameProceedings - 2023 IEEE Winter Conference on Applications of Computer Vision, WACV 2023

Conference

Conference23rd IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2023
Country/TerritoryUnited States
CityWaikoloa
Period3/01/237/01/23

Keywords

  • Algorithms: Image recognition and understanding (object detection, categorization, segmentation)
  • Machine learning architectures
  • and algorithms (including transfer, low-shot, semi-, self-, and un-supervised learning)
  • formulations

Fingerprint

Dive into the research topics of 'TransVLAD: Multi-Scale Attention-Based Global Descriptors for Visual Geo-Localization'. Together they form a unique fingerprint.

Cite this