[目的/意义]揭示社交媒体环境下识别谣言过程中的关键要素和谣言识别机制,识别突发公共卫生事件
中的谣言微博,研究及评估影响谣言识别的重要特征,有助于准确识别网络谣言、维护健康的网络生态环境。[方
法/过程]文章抽取谣言微博的用户特征、时间特征、微博文本结构特征、文本语义特征和微博传播特征,结合
MAIN理论模型,采用二元逻辑回归方法从信息内容、信息模态、信息源角度对谣言的影响因素深入研究,利用神
经网络模型提取文本语义特征,构建融合文本语义特征的多特征谣言识别模型,并通过XGBoost算法计算不同特
征在谣言识别中的重要性。[结果/结论]正向评论情感度、用户发布微博数、用户影响力越大,则是谣言的可能性
越小。谣言识别模型的准确率达到0.984,其中,文本语义特征的重要性最高。
[Purpose/significance] This study aims to explore the significance of factors on the rumor identification in
the social media environment and identify rumors in public health emergencies. We also evaluated the important features that affect the identification of rumors to help the cyber security department accurately identify rumors and maintain a healthy network ecological environment. [Method/process] We extracted user features, time features, structure
features, text semantic features and propagation features in microblog entries. We combined with the MAIN theoretical
models, and used binary logistic regression method to deeply research the influence factors of rumors from the perspective of modality, information content, information sources. We built a multi-feature based rumor identification model
that integrated the semantic feature extracted by neural network model. XGBoost algorithm was used to calculate the importance of different features in rumor identification. [Result/conclusion] The higher the positive emotional value of
comment, the number of microblog entries posted by users, and greater the influence of users, the lower the possibility
that the microblog entry is a rumor. The value of the accuracy of rumor recognition model is 0.984. The semantic features of text are the most important.