理论探讨

形式概念分析视角下融合多维特征的突破性论文识别研究——基于FCA-SHAP可解释体系

  • 李强 ,
  • 顾晓婷 ,
  • 钱智勇 ,
  • 姜衍
展开
  • (南通大学图书馆江苏226019)
李强,男,1979年生,南通大学图书馆副研究馆员。 顾晓婷,女,2001年生,南通大学图书馆硕士研究生。 钱智勇,男,1966年生,南通大学图书馆研究馆员。 姜衍,男,1981年生,南通大学图书馆副研究员。

网络出版日期: 2026-05-14

基金资助

本文系国家社会科学基金一般项目“基于关联数据的古代辞书知识组织与应用研究”(批准号:21BTQ094)、江苏省高等教育教改研究课题“双千计划背景下古籍数智保护微专业建设研究”(项目编号:2025JGYB535)和江苏省高校图工委教改研究课题重点项目“高校‘未来学习中心’建设与实践研究”(项目编号:2024JTZD10)的阶段性成果之一。

Research on the Identification of Breakthrough Papers by Integrating Multi-dimensional Features from the Perspective of Formal Concept Analysis: Based on the FCA-SHAP Explainable System

  • Li Qiang ,
  • Gu Xiaoting ,
  • Qian Zhiyong ,
  • Jiang Yan
Expand
  • (Nantong University Library, Jiangsu, 226019)

Online published: 2026-05-14

摘要

[目的/意义]针对当前突破性论文早期识别中特征维度单一、解释性不足等问题,构建可解释性机器学习方法提升识别精准度与逻辑透明度,为科研管理与创新布局提供方法论支撑。[方法/过程]首先,从突破性论文的内涵出发,在传统特征的基础上引入创新属性测度,利用大语言Prompt诱导获取知识创新向量,构建突破性论文多维特征体系;其次,基于形式概念分析理论构建信息表达系统,结合统计相关性分析及FCA属性约简算法筛选核心特征,采用多种机器学习分类器进行模型识别效果预测;最后,构建基于概念格与SHAP分析的双层解释框架,形成从筛选规则到预测验证的可视化解释链。[结果/结论]XGBoost模型的F1值达0.952,显著优于传统方法;双层解释体系明确了高创新属性与强知识关联的特征组合规则,量化了单特征对预测结果的贡献度。

本文引用格式

李强 , 顾晓婷 , 钱智勇 , 姜衍 . 形式概念分析视角下融合多维特征的突破性论文识别研究——基于FCA-SHAP可解释体系[J]. 情报资料工作, 2026 , 47(3) : 24 -32 . DOI: 10.12154/j.qbzlgz.2026.03.003

Abstract

[Purpose/significance] Aiming at the problems of single feature dimension and insufficient interpretability in the early identification of current breakthrough papers, this study constructs an interpretable machine learning meth⁃od to improve the identification accuracy and logical transparency, providing methodological support for scientific re⁃search management and innovation layout. [Method/process] Firstly, starting from the connotation of breakthrough pa⁃pers, on the basis of traditional features, innovation attribute measurement is introduced, knowledge innovation vectors are obtained by inducing large language Prompts, and a multi-dimensional feature system of breakthrough papers is constructed; secondly, an information expression system is built based on formal concept analysis (FCA), core features are screened by combining statistical correlation analysis and FCA attribute reduction algorithm, and various machine learning classifiers are used to predict the model identification effect; finally, a two-layer interpretation framework based on FCA concept lattice and SHAP analysis is constructed to form a visual interpretation chain from screening rules to prediction verification. [Result/conclusion] XGBoost model has an F1 value of 0.952 on multi-disciplinary da⁃tasets, which is significantly better than traditional methods; the two-layer interpretation system clarifies the feature combination rules of high innovation attributes and strong knowledge correlation, and quantifies the contribution of sin⁃gle features to the prediction results.
文章导航

/