Research on Large Language Model for Automatic Part-of-speech Tagging Based on Classics and Cross-language

  • Liu Yang ,
  • Xu Qiankun ,
  • Liu Chang ,
  • Wang Dongbo
Expand
  • (College of Information Management, Nanjing Agricultural University, Jiangsu,210095)

Online published: 2025-03-19

Abstract

[Purpose/significance] The instruction following, chain of thought and reasoning ability of large language models provide new opportunities for the automatic part-of-speech tagging task of ancient texts, which is conducive to the promotion of the research work on intelligent information processing of ancient texts. [Method/process] In this study, based on a manually calibrated part-of-speech tagging corpus of the Twenty-four Histories, we apply the LoRA method to efficiently supervise the fine-tuning of mainstream Chinese LLMs, and test the zero-shot and one-shot learn⁃ing ability of the SFT model in order to compare its performance in word segmentation and part-of-speech tagging of ancient texts. [Result/conclusion] The test found that the fine-tuned Xunzi-Baichuan model has the best overall per⁃formance, with F1 scores of 92.293% and 85.75% for word segmentation and part-of-speech tagging of ancient texts,respectively, while the F1 scores for word segmentation and part-of-speech tagging of modern Chinese are 91.993%
and 86.344%.

Cite this article

Liu Yang , Xu Qiankun , Liu Chang , Wang Dongbo . Research on Large Language Model for Automatic Part-of-speech Tagging Based on Classics and Cross-language[J]. Information and Documentation Services, 2025 , 46(2) : 82 -90 . DOI: 10.12154/j.qbzlgz.2025.02.009

Outlines

/