Attention-based pyramid aggregation network for visual place recognition
| dc.contributor.author | Zhu, Yingying | en |
| dc.contributor.author | Xie, Lingxi | en |
| dc.contributor.author | Wang, Jiong | en |
| dc.contributor.author | Zheng, Liang | en |
| dc.date.accessioned | 2025-02-11T01:04:19Z | |
| dc.date.available | 2025-02-11T01:04:19Z | |
| dc.date.issued | 2018-10-15 | en |
| dc.description.abstract | Visual place recognition is challenging in the urban environment and is usually viewed as a large scale image retrieval task. The intrinsic challenges in place recognition exist that the confusing objects such as cars and trees frequently occur in the complex urban scene, and buildings with repetitive structures may cause over-counting and the burstiness problem degrading the image representations. To address these problems, we present an Attention-based Pyramid Aggregation Network (APANet), which is trained in an end-to-end manner for place recognition. One main component of APANet, the spatial pyramid pooling, can effectively encode the multi-size buildings containing geo-information. The other one, the attention block, is adopted as a region evaluator for suppressing the confusing regional features while highlighting the discriminative ones. When testing, we further propose a simple yet effective PCA power whitening strategy, which significantly improves the widely used PCA whitening by reasonably limiting the impact of over-counting. Experimental evaluations demonstrate that the proposed APANet outperforms the state-of-the-art methods on two place recognition benchmarks, and generalizes well on standard image retrieval datasets. | en |
| dc.description.sponsorship | This work was supported by: (i) National Natural Science Foundation of China (Grant No. 61602314); (ii) Natural Science Foundation of Guangdong Province of China (Grant No. 2016A030313043); (iii) Fundamental Research Project in the Science and Technology Plan of Shenzhen (Grant No. JCYJ20160331114551175). We would also like to thank Relja Arandjelović and Akihiko Torii for providing data, codes, and sharing insights, and Jie Lin for insightful discussions. | en |
| dc.description.status | true | en |
| dc.format.extent | 9 | en |
| dc.identifier.isbn | 9781450356657 | en |
| dc.identifier.other | researchoutputwizard:u3102795xPUB197 | en |
| dc.identifier.other | Scopus:85058240859 | en |
| dc.identifier.other | WOS:WOS:000509665700012 | en |
| dc.identifier.uri | https://dspace-test.anu.edu.au/handle/1885/733712714 | |
| dc.identifier.url | http://www.scopus.com/inward/record.url?scp=85058240859&partnerID=8YFLogxK | en |
| dc.language.iso | English | en |
| dc.relation.ispartofseries | MM 2018 - Proceedings of the 2018 ACM Multimedia Conference | en |
| dc.rights | Publisher Copyright: © 2018 Association for Computing Machinery. | en |
| dc.subject | Attention mechanism | en |
| dc.subject | Content-based image retrieval | en |
| dc.subject | Convolutional neural network | en |
| dc.subject | Place recognition | en |
| dc.title | Attention-based pyramid aggregation network for visual place recognition | en |
| dc.type | Conference contribution | en |
| local.bibliographicCitation.lastpage | 107 | en |
| local.bibliographicCitation.startpage | 99 | en |
| local.contributor.affiliation | Zhu, Yingying; Shenzhen University | en |
| local.contributor.affiliation | Xie, Lingxi; Johns Hopkins University | en |
| local.contributor.affiliation | Wang, Jiong; Shenzhen University | en |
| local.contributor.affiliation | Zheng, Liang; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.identifier.doi | 10.1145/3240508.3240525 | en |
| local.identifier.pure | 2784a2cf-ad3b-44f0-bf0f-56d3552d364c | en |
| local.type.status | Published | en |