[目的/意义]从古文到现代文的机器翻译过程中,由于古文与现代文之间在词汇构成、句法以及词类活用
等方面的显著差异,并且缺少公开的古文分词数据,使得机器翻译系统对古文的理解和处理能力存在偏差,一定
程度上影响了翻译的质量。[方法/过程]文章提出无监督词库构建的方法,在UniLM模型的基础上,分别与BERT、
RoBERTa、RoFormer和RoFormerV2预训练模型相结合并对模型进行微调,借助UniLM模型融合古文领域知识特
征将源语言和目标语言之间的语言关系生成中间的语言表示,利用预训练模型学习上下文相关的语言表示,增加
语义之间的关联性,从而提升古现机器翻译的性能。[结果/结论]实验结果表明,融合古文领域知识特征的古文机
器翻译在BERT、RoBERTa、RoFormer和RoFormerV2预训练模型上的BLEU值分别提高了0.27到1.12,证明了提
出方法的有效性。
[Purpose/significance] In the process of machine translation from ancient Chinese to modern Chinese, due
to the significant differences in vocabulary composition, syntax and flexible use of parts of speech between ancient Chi⁃
nese and modern Chinese, and the lack of open word segmentation data of ancient Chinese, the understanding and pro⁃
cessing ability of machine translation system is biased, which affects the translation quality to some extent. [Method/
process] Firstly, this paper puts forward an unsupervised thesaurus construction method. Based on UniLM model, it is
combined with BERT, RoBERTa, RoFormer and RoFormerV2 pre-training models respectively and fine-tuned the
model. With the help of UniLM model, the language relationship between the source language and the target language
is generated into an intermediate language representation, and the pre-training model is used to learn the context-relat⁃
ed language representation, so as to increase the relevance between semantics, thus improving the machine translation
of ancient and modern times. [Result/conclusion] The experimental results show that the BLEU value of machine
translation of ancient Chinese prose, which integrates the knowledge characteristics of ancient Chinese prose, is in⁃
creased by 0.27 to 1.12 on BERT, RoBERTa, RoFormer and RoFormerV2 pre- training models respectively, which
proves the effectiveness of the proposed method.