论文部分内容阅读
将信息增益和加权log似然比特征选择方法应用于音子配列学语种识别系统中进行特征降维。在美国国家标准技术研究院2009年语种识别评测数据集上进行实验,分别使用信息增益和加权log似然比准则以及传统的互信息,χ~2统计量方法对数量巨大的N-gram进行特征选择,从中选出最具有鉴别性的部分组成特征向量,并用分类器进行分类。结果显示,当根据信息增益和加权log似然比准则选取一定数量的特征时,系统性能与使用全部特征的基线系统相比略好;当选取的特征数量很少时,信息增益和加权log似然比方法的性能要优于传统的互信息和χ~2统计量方法。实验表明,在音子配列学语种识别系统中,信息增益和加权log似然比方法均可以有效地去除冗余信息,降低特征向量的维数,并且能使系统性能得到一定的提高。
The information gain and weighted log likelihood ratio feature selection method is applied to the phonon collocation language recognition system for feature dimension reduction. Experiments were carried out on the 2009 National Institute of Standards and Technology Chinese Language Recognition Evaluation dataset, using the information gain and weighted log likelihood ratio criteria and the traditional mutual information, χ ~ 2 statistical methods to characterize a large number of N-grams Select, select the most discriminating part of the composition of eigenvectors, and classified by the classifier. The results show that when a certain number of features are selected according to the information gain and the weighted log likelihood ratio criterion, the system performance is slightly better than the baseline system using all the features. When the number of selected features is small, the information gain and the weighted log However, the performance of the method is superior to the traditional method of mutual information and χ ~ 2 statistics. Experiments show that both information gain and weighted log likelihood ratio method can effectively remove redundant information and reduce the dimensionality of feature vector in the phonetic collocations language recognition system, and can improve the system performance to a certain extent.