信息技术

基于典籍跨语言的自动词性标注大语言模型研究

  • 刘洋 ,
  • 许乾坤 ,
  • 刘畅 ,
  • 王东波
展开
  • (南京农业大学信息管理学院江苏210095)
刘洋,男,1998年生,南京农业大学信息管理学院博士研究生。 许乾坤,男,1996年生,南京农业大学信息管理学院博士研究生。 刘畅,男,1998年生,南京农业大学信息管理学院博士研究生。 王东波,男,1981年生,南京农业大学信息管理学院教授,博士生导师。

网络出版日期: 2025-03-19

基金资助

本文系国家社会科学基金重大项目“中国古代典籍跨语言知识库构建及应用研究”(批准号:21&ZD331)的研究成果之一。

Research on Large Language Model for Automatic Part-of-speech Tagging Based on Classics and Cross-language

  • Liu Yang ,
  • Xu Qiankun ,
  • Liu Chang ,
  • Wang Dongbo
Expand
  • (College of Information Management, Nanjing Agricultural University, Jiangsu,210095)

Online published: 2025-03-19

摘要

[目的/意义]大语言模型的指令遵循、思维链及推理能力为古籍文本的自动词性标注任务提供了新的契机,有利于促进古籍智能信息处理研究工作的开展。[方法/过程]以人工校验的《二十四史》古现词性标注语料为基础,运用LoRA方法对主流中文大语言模型进行高效监督微调,并测试SFT模型的零样本(Zero-shot)和单样本(One-shot)学习能力,以比较其在古现文本分词与词性标注上的性能。[结果/结论]经过微调后的Xunzi-Baichuan模型整体表现最优,古文本的分词和词性标注F1得分分别为92.293%和85.75%,现代汉语的分词和词性标注F1得分为91.993%和86.344%。

本文引用格式

刘洋 , 许乾坤 , 刘畅 , 王东波 . 基于典籍跨语言的自动词性标注大语言模型研究[J]. 情报资料工作, 2025 , 46(2) : 82 -90 . DOI: 10.12154/j.qbzlgz.2025.02.009

Abstract

[Purpose/significance] The instruction following, chain of thought and reasoning ability of large language models provide new opportunities for the automatic part-of-speech tagging task of ancient texts, which is conducive to the promotion of the research work on intelligent information processing of ancient texts. [Method/process] In this study, based on a manually calibrated part-of-speech tagging corpus of the Twenty-four Histories, we apply the LoRA method to efficiently supervise the fine-tuning of mainstream Chinese LLMs, and test the zero-shot and one-shot learn⁃ing ability of the SFT model in order to compare its performance in word segmentation and part-of-speech tagging of ancient texts. [Result/conclusion] The test found that the fine-tuned Xunzi-Baichuan model has the best overall per⁃formance, with F1 scores of 92.293% and 85.75% for word segmentation and part-of-speech tagging of ancient texts,respectively, while the F1 scores for word segmentation and part-of-speech tagging of modern Chinese are 91.993%
and 86.344%.
文章导航

/