信息技术

基于GenAI数据增强的突发事件毒性内容检测研究

  • 邓胜利 ,
  • 刘李毅 ,
  • 朱秋雨 ,
  • 程麟淇
展开
  • (武汉大学信息管理学院湖北430072)
邓胜利,男,1979年生,武汉大学信息管理学院教授。 刘李毅,男,1998年生,武汉大学信息管理学院博士研究生(通讯作者)。 朱秋雨,女,2000年生,武汉大学信息管理学院博士研究生。 程麟淇,男,1999年生,武汉大学信息管理学院博士研究生。

网络出版日期: 2026-01-20

基金资助

本文系国家社会科学基金重大项目“信息资源管理学科研究方法知识库构建及其应用研究”(批准号:23&ZD229)的研究成果之一。

Research on Toxic Content Detection in Emergency Events Based on GenAI Data Augmentation

  • Deng Shengli ,
  • Liu Liyi ,
  • Zhu Qiuyu ,
  • Cheng Linqi
Expand
  • (School of Information Management, Wuhan University, Hubei, 430072)

Online published: 2026-01-20

摘要

[目的/意义]突发事件中,社交媒体毒性内容因舆情传播放大效应,对网络环境与应急管理构成严峻挑战。传统检测方法存在分类数据不平衡、检测精度低等问题,亟需高效解决方案。[方法/过程]构建突发事件场景下的微博毒性内容专用数据集,采用生成式人工智能(GenAI)技术,通过少样本提示学习生成语义等效的伪毒性内容以平衡样本。进而提出融合MACBert预训练模型与全局注意力机制的MACBert-Att模型,强化对领域术语与情绪化表达的语义捕捉能力。[结果/结论]经GenAI数据增强后,MACBert-Att模型F1值达0.95,较基线Bert模型提升15%,显著优于SMOTE等传统增强方法,验证了GenAI语义级数据增强与模型架构的协同有效性。

本文引用格式

邓胜利 , 刘李毅 , 朱秋雨 , 程麟淇 . 基于GenAI数据增强的突发事件毒性内容检测研究[J]. 情报资料工作, 2026 , 47(1) : 86 -93 . DOI: 10.12154/j.qbzlgz.2026.01.009

Abstract

[Purpose/significance] In the context of sudden emergency events, content on social media poses severe challenges to online environments and emergency management due to the amplification effect of public opinion dissemi⁃nation. Traditional detection methods suffer from issues such as imbalanced classification data and low detection accu⁃racy, necessitating efficient solutions. [Method/process] This study constructs a dedicated dataset of content from Wei⁃bo comments in sudden emergency events. Generative Artificial Intelligence (GenAI) technology is employed to gener⁃ate semantically equivalent pseudo-toxic content through few-shot prompt learning, balancing the sample distribution.Furthermore, the MACBert-Att model is proposed by integrating the MACBert pre-trained model with a global atten⁃tion mechanism, enhancing the semantic capture capability for domain-specific terms and emotional expressions. [Re⁃sult/conclusion] Experiments demonstrate that with GenAI data augmentation, the MACBert-Att model achieves an F1 score of 0.95, representing a 15% improvement over the baseline Bert model and significantly outperforming tradi⁃tional augmentation methods like SMOTE. This validates the collaborative effectiveness of GenAI-based semantic-lev⁃el data augmentation and the model architecture.
文章导航

/