[Purpose/significance] Classifying research topics into their respective disciplines is fundamental for inter⁃disciplinary research, such as measuring interdisciplinary integration and identifying cross-disciplinary themes. Only by determining the disciplinary category of each research topic can we assess whether it represents an interdisciplinary intersection. [Method/process] This study proposes a framework for classifying research topics into disciplines using large language models. Building on base models such as Llama3-8B-Instruct, Qwen2.5-7B-Instruct, and DeepSeek-R1-Distill-Qwen-7B, a two-stage optimization strategy of "domain-adaptive pretraining + supervised fine-tuning" was implemented. A total of 116192 academic papers were used for domain-adaptive pretraining to enhance the model′s semantic understanding of scientific literature. Author keywords were used to represent research topics, and a manual⁃ly annotated dataset of 126919 "research topic-disciplinary label" pairs was employed for supervised fine-tuning to op⁃timize the model′s classification performance. [Result/conclusion] While large language models possess zero-shot dis⁃ciplinary classification capabilities, their precision and F1-scores remain below 50% when relying solely on prompt de⁃sign, which is insufficient for practical applications. In contrast, the proposed framework achieves a precision of 93.61% and an F1-score of 83.09%, significantly improving the accuracy of disciplinary classification for research top⁃ics.
Huo Chaoguang
,
Wang Xiaoyu
,
Yan Peng
. Research on the Classification Method of Disciplines for Research Topics Based on Large Models[J]. Information and Documentation Services, 2026
, 47(2)
: 69
-76
.
DOI: 10.12154/j.qbzlgz.2026.02.008