Research on Named Entity Recognition of the Documents of History of the Communist Partyof China Based on Deep Learning

  • Cao Shujin ,
  • Yue Wenyu
Expand
  • School of Information Management,Sun-Yat-Sen University, Guangzhou,510006

Online published: 2022-09-14

Abstract

[Purpose/significance] The naming entity recognition of the Communist Party of China history documents based on deep learning is helpful to explore and excavate the value of the Communist Party of China history resources. It plays a fundamental role in the construction of professional terminology database, knowledge map and knowledge question answering system in the field of the Communist Party of China history, and provides basic support for further intelligent digital humanities research of the Communist Party of China history. [Method/process] In this study, a string matching algorithm based on Trie-Tree is used to complete the batch annotation task of the experimental corpus. The XLNet (Generalized Autoregressive Pretraining for Language Understanding, XLNet) pre-training model is used to embed the mainstream BiLSTM-CRF model to construct a named entity recognition model based on XLNet-BiLSTMCRF for the history documents of the Communist Party of China . [Result/conclusion] The model performs well in named entity recognition with a harmonic mean F value of 0.9535, which is higher than the deep learning models such as BiLSTM-CRF, BERT-BiLSTM-CRF, BERT-wwm-ext-BiLSTM-CRF, and XLNet-CRF . The study shows that the method proposed in this paper is feasible and effective for the unstructured text mining work of the Communist Par? ty of China history. 

Cite this article

Cao Shujin , Yue Wenyu . Research on Named Entity Recognition of the Documents of History of the Communist Partyof China Based on Deep Learning[J]. Information and Documentation Services, 2022 , 43(5) : 81 -88 . DOI: 10.12154/j.qbzlgz.2022.05.009

Outlines

/