[Purpose /significance]It is important for us to maximize use of digital culture and tourism resources by
Image-TextCross Modal Retrieval(IT-CMR).Related methods used in the field of digital culture and tourism resources
have the challenges of long text, limited memory and some missed data. In order to address those problems, we pro?
posed a new method ofIT-CMR using Transformer and MobileNet V3 models for digital culture and tourism resources.
[Method /process]Two Layer Multi- group Transformers(TLMT) model based on attention network is proposed to
learn the complemental text features from the title, main text and comments. Local fine- grained image features are
learned using Fast R-CNN and MobileNet V3 models. Multiple linear regression model is proposed to synthesize the
missed data in the shared sub- space. Bi- directional triplet loss function for searching images by text and searching
text by image is constructed to learn the parameters of network.[Result /conclusion] Extensive experimental results on
standard benchmark Flickr30k,our own dataset CulTour-Sha, and two datasets Flickr30k-1 and CulTour-Sha-1 in?
cluding some missed data demonstrate that: our method has better recall, need less memory space and has faster com?
puting speed than several state-of-art methods of ITCMR.
Gao Yunmei
. Digital Cultural Travel-oriented Image and Text Cross-modal Retrieval Method[J]. Information and Documentation Services, 2022
, 43(1)
: 71
-80
.
DOI: 10.12154/j.qbzlgz.2022.01.006