信息技术

基于大模型的研究主题所属学科分类方法研究

  • 霍朝光 ,
  • 王晓玉 ,
  • 燕鹏
展开
  • 1中国人民大学信息资源管理学院北京100872; 2中国人民大学数字人文研究院北京100872)
霍朝光,男,1990年生,中国人民大学信息资源管理学院副教授。 王晓玉,女,1997年生,中国人民大学信息资源管理学院硕士研究生(通讯作者)。 燕鹏,男,1991年生,中国人民大学信息资源管理学院硕士研究生。

网络出版日期: 2026-03-16

基金资助

本文系国家自然科学基金面上项目“基于图机器学习的学科交叉主题识别与预测研究”(批准号:72374202)、中国人民大学科学研究基金项目(中央高校基本科研业务费专项资金、国家治理大数据和人工智能创新平台经费资助)的研究成果。

Research on the Classification Method of Disciplines for Research Topics Based on Large Models

  • Huo Chaoguang ,
  • Wang Xiaoyu ,
  • Yan Peng
Expand
  • 1School of Information Resource Management, Renmin University of China, Beijing, 100872;2School of Digital Humanities, Renmin University of China, Beijing, 100872)

Online published: 2026-03-16

摘要

[目的/意义]研究主题所属学科分类是学科交叉测度以及学科交叉主题识别等学科交叉以及跨学科研究的基础,只有确定每个研究主题所属的学科门类,才能判断其是否为学科交叉点。[方法/过程]构建基于大模型的研究主题学科分类框架,在Llama3-8B-Instruct、Qwen2.5-7B-Instruct和DeepSeek-R1-Distill-Qwen-7B三个基座模型基础上,采用“领域自适应预训练+监督微调”双阶段优化策略,利用116192篇学术论文进行领域预训练,以增强模型对科学文献的语义理解能力,以作者自标注关键词表征研究主题,通过人工标注的126919条“研究主题—学科标签”数据集进行监督微调,以优化模型对研究主题的学科分类能力。[结果/结论]虽然大语言模型具有零样本学科分类能力,但是仅仅靠设计Prompt使用大模型的精确率和F1值均低于50%,难以满足实际需要;而构建的研究主题所属学科分类框架精确率达93.61%、F1值达83.09%,显著提升了研究主题所属学科分类效果。

本文引用格式

霍朝光 , 王晓玉 , 燕鹏 . 基于大模型的研究主题所属学科分类方法研究[J]. 情报资料工作, 2026 , 47(2) : 69 -76 . DOI: 10.12154/j.qbzlgz.2026.02.008

Abstract

[Purpose/significance] Classifying research topics into their respective disciplines is fundamental for inter⁃disciplinary research, such as measuring interdisciplinary integration and identifying cross-disciplinary themes. Only by determining the disciplinary category of each research topic can we assess whether it represents an interdisciplinary intersection. [Method/process] This study proposes a framework for classifying research topics into disciplines using large language models. Building on base models such as Llama3-8B-Instruct, Qwen2.5-7B-Instruct, and DeepSeek-R1-Distill-Qwen-7B, a two-stage optimization strategy of "domain-adaptive pretraining + supervised fine-tuning" was implemented. A total of 116192 academic papers were used for domain-adaptive pretraining to enhance the model′s semantic understanding of scientific literature. Author keywords were used to represent research topics, and a manual⁃ly annotated dataset of 126919 "research topic-disciplinary label" pairs was employed for supervised fine-tuning to op⁃timize the model′s classification performance. [Result/conclusion] While large language models possess zero-shot dis⁃ciplinary classification capabilities, their precision and F1-scores remain below 50% when relying solely on prompt de⁃sign, which is insufficient for practical applications. In contrast, the proposed framework achieves a precision of 93.61% and an F1-score of 83.09%, significantly improving the accuracy of disciplinary classification for research top⁃ics.
文章导航

/