论文部分内容阅读
针对网络评论挖掘中的产品特征抽取准确度不高、人工参与较多和难以处理口语化表述等问题,提出一种基于潜在狄利特雷分布模型的产品特征抽取方法。该方法首先应用中文分词工具对网络评论信息进行分词和词性标注,得到最初的产品特征名词集合;然后采用潜在狄利特雷分布文本训练模型筛选出候选产品特征词集合,进而通过同义词词林拓展和过滤规则得到最终的产品特征集合。以京东网上的相机和手机评论数据为例,通过实验对比分析验证了所提方法的有效性。
Aiming at the problems such as low accuracy of product feature extraction, large amount of manual participation, and difficulty in processing colloquial expression in network comment mining, a product feature extraction method based on latent Dilithre distribution model was proposed. The method first uses the Chinese word segmentation tool to segment the word-by-word and part-of-speech tagging of the online commentary information, and obtains the initial set of product feature nouns. Then, the potential Dilithre distribution text training model is used to filter out the candidate product feature word set, and further expands through the synonym word forest. And filtering rules to get the final product feature set. Taking Jingdong Online’s camera and mobile phone commentary data as an example, the effectiveness of the proposed method is verified through experimental comparative analysis.