信息技术

融合不同语义知识的中国古代典籍机器翻译研究

  • 吴梦成 ,
  • 林立涛 ,
  • 吴 娜 ,
  • 许乾坤 ,
  • 王东波
展开
  • 1 南京农业大学信息管理学院 江苏 210095; 

    2 南京农业大学人文与社会计算研究中心 江苏 210095; 

    3 南京农业大学领域知识关联研究中心 江苏 210095; 

    4 南京大学信息管理学院 江苏 210023

吴梦成,男,1997年生,南京农业大学信息管理学院博士研究生。 林立涛,男,1997年生,南京大学信息管理学院博士研究生。 吴 娜,女,1999年生,南京农业大学信息管理学院硕士研究生。 许乾坤,男,1996年生,南京农业大学信息管理学院博士研究生。 王东波,男,1981年生,南京农业大学信息管理学院教授,博士生导师(通讯作者)

网络出版日期: 2024-03-14

基金资助

本文系国家社会科学基金重大项目“中国古代典籍跨语言知识库构建及应用研究”(批准号:21&ZD331)的研究成果之一。

Research on Machine Translation of Ancient Chinese Classics by Integrating Different Semantic Knowledge

  • Wu Mengcheng ,
  • Lin Litao ,
  • Wu Na ,
  • Xu Qiankun ,
  • Wang Dongbo
Expand
  • 1 College of Information Management, Nanjing Agricultural University, Jiangsu, 210095; 

    2 Research Center for Humanities and Social Computing, Nanjing Agricultural University, Jiangsu, 210095; 

    3 Research Center for Correlation of Domain Knowledge, Nanjing Agricultural University, Jiangsu, 210095; 

    4 School of Information Management, Nanjing University, Jiangsu, 210023


Online published: 2024-03-14

摘要

[目的/意义]文章旨在探究将不同语义知识融入机器翻译模型能否增强机器翻译的效果以及何种语义知 识的作用更为显著,以助力机器翻译研究与中华优秀传统文化的传承与传播。[方法/过程]研究选取了30万对精 加工的《二十四史》“古代汉语-现代汉语”平行语料作为实验数据,基于神经机器翻译OpenNMT模型,通过三种不 同的特征融合方法,将词边界知识、词性知识、实体知识和依存句法知识分别融入机器翻译模型的训练过程中。 [结果/结论]不同语义知识与模型的融合对典籍翻译效果有不同的影响,词边界知识、词性知识、实体知识对机器 翻译任务有一定的贡献且实体知识的贡献最大,依存句法知识无明显作用。

本文引用格式

吴梦成 , 林立涛 , 吴 娜 , 许乾坤 , 王东波 . 融合不同语义知识的中国古代典籍机器翻译研究[J]. 情报资料工作, 2024 , 45(2) : 97 -104 . DOI: 10.12154/j.qbzlgz.2024.02.011

Abstract

[Purpose/significance] This article aims to explore whether integrating different semantic knowledge into machine translation models can enhance the effectiveness of machine translation and which type of semantic knowl⁃ edge plays a more significant role. The purpose is to support the research in machine translation and the inheritance and dissemination of Chinese excellent traditional culture. [Method/process] The study selected 300,000 pairs of me⁃ ticulously processed "Ancient Chinese-Modern Chinese" parallel corpora from the "Twenty-Four Histories" as experi⁃ mental data. Based on the neural machine translation model OpenNMT, it integrated word boundary knowledge, partof- speech knowledge, entity knowledge, and dependency syntax knowledge into the training process of the machine translation model through three different feature fusion methods. [Result/conclusion] The integration of different se⁃ mantic knowledge with the model has varying impacts on the translation effectiveness of classical texts. Word boundary knowledge, part- of- speech knowledge, and entity knowledge contribute to the machine translation task, with entity knowledge making the largest contribution, while the role of dependency syntax knowledge has no obvious effect.
文章导航

/