[目的/意义]构建适用于开放数据环境下的中文政策文本分析情感词典,对精准把握政府行为施政理念具有重要价值。[方法/过程]将政策文本中体现情感强度的词汇定义为倾向词,利用其结构特征和语义关联性构建倾向词典。首先,依据领域专家解读意见抽取种子词并结合点互信息算法进行词典在线扩充。其次,基于形式概念分析理论定义并量化政策文本主题内涵,将政策主题间的层次关系映射到词汇间的语义关系,筛选具有主题相似关系的同义倾向词。最后,采用可信度与有效度方法进行实证检验。[结果/结论]倾向词典在政策文本情感识别任务中具有较高的准确率与召回率,适用于大规模细粒度政策文本分析,该方法将为政策信息学研究提供切实可行的量化工具。
[Purpose/significance] It is of great value to grasp the concept of government behavior and governance ac⁃curately, based on the dictionary method to construct a fine grain analysis of Chinese policy texts suitable for the open data environment. [Method/process] The words that reflect emotional intensity in policy texts are defined as tendency words, and their structural characteristics and semantic relevance are used to construct a tendency dictionary. First, the seed words are extracted according to the interpretation opinions of domain experts and combined with the point mutual information algorithm to expand the dictionary online. Secondly, based on the theory of formal concept analysis, the top⁃ic connotation of policy texts is defined and quantified, the hierarchical relationship between policy topics is mapped to the semantic relationship between words, and the synonymous tendency words with topic similarity are screened. Final⁃ly, the credibility and validity methods are used for empirical testing. [Result/conclusion] The propensity dictionary
has high accuracy and recall rate in the task of policy text emotion recognition, which is suitable for large-scale fine grain policy text analysis, and provides a reliable and novel quantitative tool for policy research and decision-making.