Digital Cultural Travel-oriented Image and Text Cross-modal Retrieval Method

  • Gao Yunmei
Expand
  •  Library of Changshu Institute of Technology, Jiangsu, 215500

Online published: 2022-01-18

Abstract

[Purpose /significance]It is important for us to maximize use of digital culture and tourism resources by Image-TextCross Modal Retrieval(IT-CMR).Related methods used in the field of digital culture and tourism resources have the challenges of long text, limited memory and some missed data. In order to address those problems, we pro? posed a new method ofIT-CMR using Transformer and MobileNet V3 models for digital culture and tourism resources. [Method /process]Two Layer Multi- group Transformers(TLMT) model based on attention network is proposed to learn the complemental text features from the title, main text and comments. Local fine- grained image features are learned using Fast R-CNN and MobileNet V3 models. Multiple linear regression model is proposed to synthesize the missed data in the shared sub- space. Bi- directional triplet loss function for searching images by text and searching text by image is constructed to learn the parameters of network.[Result /conclusion] Extensive experimental results on standard benchmark Flickr30k,our own dataset CulTour-Sha, and two datasets Flickr30k-1 and CulTour-Sha-1 in? cluding some missed data demonstrate that: our method has better recall, need less memory space and has faster com? puting speed than several state-of-art methods of ITCMR.

Cite this article

Gao Yunmei . Digital Cultural Travel-oriented Image and Text Cross-modal Retrieval Method[J]. Information and Documentation Services, 2022 , 43(1) : 71 -80 . DOI: 10.12154/j.qbzlgz.2022.01.006

Outlines

/