理论探讨

基于多元数据融合的科学文献主题识别研究

  • 邱均平 ,
  • 孙月瑞 ,
  • 周贞云
展开
  • 1 杭州电子科技大学中国科教评价研究院 浙江 310018;   2 杭州电子科技大学管理学院 浙江 310018;  3 杭州电子科技大学数据科学与信息计量研究院 浙江 310018
邱均平,男,1947年生,杭州电子科技大学中国科教评价研究院博士生导师,资深教授。 孙月瑞,男,1997年生,杭州电子科技大学管理学院硕士研究生在读。 周贞云,男,1979年生,杭州电子科技大学中国科教评价研究院 博士研究生,副教授(通讯作者)。

网络出版日期: 2022-11-11

基金资助

本文系2019年国家社会科学基金重大项目“基于大数据的科教评价信息云平台构建和智能服务研究”(项目编号:19ZDA348)和2020年浙江 省软科学研究计划重点项目“创新强省背景下浙江高校科技创新竞争力评价及提升研究”(项目编号:2020C25027)的研究成果之一。

Research on the Topic Identification of Scientific Literature Based on Multivariate Data Fusion

  • Qiu Junping ,
  • Sun Yuerui ,
  • Zhou Zhenyun
Expand
  • 1 Chinese Academy of Science and Education Evaluation, Hangzhou Dianzi University, Zhejiang,310018;  2 School of Management, Hangzhou Dianzi University, Zhejiang,310018;  3 Academy of Data Science and Informatics, Hangzhou Dianzi University, Zhejiang,310018

Online published: 2022-11-11

摘要

[目的/意义]科学文献的主题识别研究是科研管理的重要内容之一,如何全面把握文献的多元数据、提升 自动文献主题识别的效果是一个值得研究的问题。[方法/过程]文献的关键词、摘要是判断文献主题的重要依据, 文章提出基于文献多元数据融合的主题识别模型,使用Word2vec模型、AP聚类及Node2vec模型表示出关键词层 的主题向量,使用LDA模型表示出摘要层的主题向量,通过多视图聚类中的SGF方法进行数据融合并识别文献 主题。[结果/结论]以不同规模的文献集为例,通过主题识别研究,验证该模型识别效果的准确性和可解释性优于 典型LDA方法、Doc-LDA模型。

本文引用格式

邱均平 , 孙月瑞 , 周贞云 . 基于多元数据融合的科学文献主题识别研究[J]. 情报资料工作, 2022 , 43(6) : 14 -20 . DOI: 10.12154/j.qbzlgz.2022.06.002

Abstract

[Purpose/significance] The research on topic identification of scientific literature is one of the important contents of scientific research management. How to comprehensively grasp the multivariate data of literature and effec? tively improve the accuracy of automatic literature topic identification is a problem worthy of research. [Method/pro? cess] Keywords and abstracts of documents are important basis for judging document topics. This paper proposes a top? ic identification model based on multi-data fusion of documents. Word2vec model, AP clustering and Node2vec model are used to represent the topic vector of the keyword layer. The topic vector of the abstract layer is represented by the LDA model, and the SGF method in the multi-view clustering method is used to perform data fusion and extract docu? ment topics. [Result/conclusion] Taking document sets of different scales as an example, through topic identification research, it is verified that the accuracy and interpretability of the recognition effect of the model are better than the typ? ical LDA method and the Doc-LDA model.
文章导航

/