信息资源

基于深度学习的中共党史文献命名实体识别研究

  • 曹树金 ,
  • 岳文玉
展开
  • 中山大学信息管理学院 广州 510006
曹树金,男,1962年生,中山大学信息管理学院教授,博士生导师。 岳文玉,女,1992年生,中山大学信息管理学院博士研究生。

网络出版日期: 2022-09-14

基金资助

本文系国家社会科学基金社科学术社团主题学术活动资助课题研究类项目“中国共产党历史知识图谱与知识索引构建研究”(项目编号: 21STA028)的研究成果之一。

Research on Named Entity Recognition of the Documents of History of the Communist Partyof China Based on Deep Learning

  • Cao Shujin ,
  • Yue Wenyu
Expand
  • School of Information Management,Sun-Yat-Sen University, Guangzhou,510006

Online published: 2022-09-14

摘要

[目的/意义]基于深度学习的中共党史文献命名实体识别,有助于探索与挖掘党史资源的价值,对于构建 中共党史领域专业术语库、知识图谱、知识问答系统等应用发挥着基础性作用,为进一步的智慧化的中共党史数 字人文研究提供基础支撑。[方法/过程]本研究采用基于Trie树的字符串匹配算法完成实验语料的批量标注任 务,利用中文XLNet(Generalized Autoregressive Pretraining for Language Understanding,XLNet)预训练模型嵌入主 流BiLSTM-CRF模型中,构建基于XLNet-BiLSTM-CRF的中共党史文献命名实体识别模型。[结果/结论]该模型 在命名实体识别中表现优异,其调和平均数F值为0.9535,高于BiLSTM-CRF、BERT-BiLSTM-CRF、BERT-wwmext-BiLSTM-CRF、XLNet-CRF等深度学习模型。研究表明本文提出的方法对于中共党史非结构化文本挖掘工 作具有可行性和有效性。

本文引用格式

曹树金 , 岳文玉 . 基于深度学习的中共党史文献命名实体识别研究[J]. 情报资料工作, 2022 , 43(5) : 81 -88 . DOI: 10.12154/j.qbzlgz.2022.05.009

Abstract

[Purpose/significance] The naming entity recognition of the Communist Party of China history documents based on deep learning is helpful to explore and excavate the value of the Communist Party of China history resources. It plays a fundamental role in the construction of professional terminology database, knowledge map and knowledge question answering system in the field of the Communist Party of China history, and provides basic support for further intelligent digital humanities research of the Communist Party of China history. [Method/process] In this study, a string matching algorithm based on Trie-Tree is used to complete the batch annotation task of the experimental corpus. The XLNet (Generalized Autoregressive Pretraining for Language Understanding, XLNet) pre-training model is used to embed the mainstream BiLSTM-CRF model to construct a named entity recognition model based on XLNet-BiLSTMCRF for the history documents of the Communist Party of China . [Result/conclusion] The model performs well in named entity recognition with a harmonic mean F value of 0.9535, which is higher than the deep learning models such as BiLSTM-CRF, BERT-BiLSTM-CRF, BERT-wwm-ext-BiLSTM-CRF, and XLNet-CRF . The study shows that the method proposed in this paper is feasible and effective for the unstructured text mining work of the Communist Par? ty of China history. 
文章导航

/