[目的/意义]基于深度学习的中共党史文献命名实体识别,有助于探索与挖掘党史资源的价值,对于构建 中共党史领域专业术语库、知识图谱、知识问答系统等应用发挥着基础性作用,为进一步的智慧化的中共党史数 字人文研究提供基础支撑。[方法/过程]本研究采用基于Trie树的字符串匹配算法完成实验语料的批量标注任 务,利用中文XLNet(Generalized Autoregressive Pretraining for Language Understanding,XLNet)预训练模型嵌入主 流BiLSTM-CRF模型中,构建基于XLNet-BiLSTM-CRF的中共党史文献命名实体识别模型。[结果/结论]该模型 在命名实体识别中表现优异,其调和平均数F值为0.9535,高于BiLSTM-CRF、BERT-BiLSTM-CRF、BERT-wwmext-BiLSTM-CRF、XLNet-CRF等深度学习模型。研究表明本文提出的方法对于中共党史非结构化文本挖掘工 作具有可行性和有效性。
[Purpose/significance] The naming entity recognition of the Communist Party of China history documents
based on deep learning is helpful to explore and excavate the value of the Communist Party of China history resources.
It plays a fundamental role in the construction of professional terminology database, knowledge map and knowledge
question answering system in the field of the Communist Party of China history, and provides basic support for further
intelligent digital humanities research of the Communist Party of China history. [Method/process] In this study, a
string matching algorithm based on Trie-Tree is used to complete the batch annotation task of the experimental corpus.
The XLNet (Generalized Autoregressive Pretraining for Language Understanding, XLNet) pre-training model is used to
embed the mainstream BiLSTM-CRF model to construct a named entity recognition model based on XLNet-BiLSTMCRF for the history documents of the Communist Party of China . [Result/conclusion] The model performs well in
named entity recognition with a harmonic mean F value of 0.9535, which is higher than the deep learning models such
as BiLSTM-CRF, BERT-BiLSTM-CRF, BERT-wwm-ext-BiLSTM-CRF, and XLNet-CRF . The study shows that
the method proposed in this paper is feasible and effective for the unstructured text mining work of the Communist Par?
ty of China history.