信息技术

面向数字文旅的图像文本跨模态检索方法

  • 高蕴梅
展开
  • 常熟理工学院图书馆 江苏 215500
高蕴梅,女,1982年生,常熟理工学院图书馆馆员。

网络出版日期: 2022-01-18

基金资助

本文系教育部人文社会科学研究青年基金项目“学术信息网络中的核心边缘结构测度研究”(项目编号:18YJC870011)、江苏高校哲学社会科 学研究项目“视频内容分析的360度教学评价研究”(项目编号:2020SJA1425)、苏州市图书馆学会2021年重点项目“视频资源跨模态智慧检索 研究”(项目编号:21-A-02)和常熟理工学院高等教育研究项目“视频内容分析的360度教学评价”(项目编号:GJ1905)的研究成果之一。

Digital Cultural Travel-oriented Image and Text Cross-modal Retrieval Method

  • Gao Yunmei
Expand
  •  Library of Changshu Institute of Technology, Jiangsu, 215500

Online published: 2022-01-18

摘要

[目的/意义]图像文本跨模态检索应用对最大化利用数字文旅资源具有重要意义。然而,数字文旅领域 的图像文本跨模态检索方法面临长文本挑战、数据缺失、内存资源有限等问题。为此,我们提出了一种新的基于 Transformers和MobileNet V3模型的数字文旅图像文本跨模态方法。[方法/过程]首先,提出了基于自注意力机制 的双层多组Transformers模型从标题、正文和评论等文本中学习具有互补性的文本特征;其次,设计了FastR-CNN 和MobileNet V3模型学习图像局部细粒度特征;最后,提出了多元线性回归方法在共享子空间补全缺失数据。构 建以图搜文和以文搜图的双向三元损失函数学习模型参数。[结果/结论]在标准数据集Flickr30k、自建数据集 CulTour-Sha和有数据缺失的数据集Flickr30k-1与CulTour-Sha-1上的大量实验结果表明,我们的方法在召回 率、内存需求和计算速度等方面优于当前几种先进的跨模态检索方法。

本文引用格式

高蕴梅 . 面向数字文旅的图像文本跨模态检索方法[J]. 情报资料工作, 2022 , 43(1) : 71 -80 . DOI: 10.12154/j.qbzlgz.2022.01.006

Abstract

[Purpose /significance]It is important for us to maximize use of digital culture and tourism resources by Image-TextCross Modal Retrieval(IT-CMR).Related methods used in the field of digital culture and tourism resources have the challenges of long text, limited memory and some missed data. In order to address those problems, we pro? posed a new method ofIT-CMR using Transformer and MobileNet V3 models for digital culture and tourism resources. [Method /process]Two Layer Multi- group Transformers(TLMT) model based on attention network is proposed to learn the complemental text features from the title, main text and comments. Local fine- grained image features are learned using Fast R-CNN and MobileNet V3 models. Multiple linear regression model is proposed to synthesize the missed data in the shared sub- space. Bi- directional triplet loss function for searching images by text and searching text by image is constructed to learn the parameters of network.[Result /conclusion] Extensive experimental results on standard benchmark Flickr30k,our own dataset CulTour-Sha, and two datasets Flickr30k-1 and CulTour-Sha-1 in? cluding some missed data demonstrate that: our method has better recall, need less memory space and has faster com? puting speed than several state-of-art methods of ITCMR.
文章导航

/