[Purpose/significance] The naming entity recognition of the Communist Party of China history documents
based on deep learning is helpful to explore and excavate the value of the Communist Party of China history resources.
It plays a fundamental role in the construction of professional terminology database, knowledge map and knowledge
question answering system in the field of the Communist Party of China history, and provides basic support for further
intelligent digital humanities research of the Communist Party of China history. [Method/process] In this study, a
string matching algorithm based on Trie-Tree is used to complete the batch annotation task of the experimental corpus.
The XLNet (Generalized Autoregressive Pretraining for Language Understanding, XLNet) pre-training model is used to
embed the mainstream BiLSTM-CRF model to construct a named entity recognition model based on XLNet-BiLSTMCRF for the history documents of the Communist Party of China . [Result/conclusion] The model performs well in
named entity recognition with a harmonic mean F value of 0.9535, which is higher than the deep learning models such
as BiLSTM-CRF, BERT-BiLSTM-CRF, BERT-wwm-ext-BiLSTM-CRF, and XLNet-CRF . The study shows that
the method proposed in this paper is feasible and effective for the unstructured text mining work of the Communist Par?
ty of China history.
Cao Shujin
,
Yue Wenyu
. Research on Named Entity Recognition of the Documents of History of the Communist
Partyof China Based on Deep Learning[J]. Information and Documentation Services, 2022
, 43(5)
: 81
-88
.
DOI: 10.12154/j.qbzlgz.2022.05.009