论文部分内容阅读
语音是一种短时平稳时频信号,因此大多数的研究者都通过分帧来提取情感特征。然而,分帧后提取的特征为局部特征,无法准确反应情感语音动态特性,故单纯采用局部特征往往无法构建鲁棒的情感识别系统。针对这个问题,先在不分帧的语音信号里通过多尺度最优小波包分解提取语句级全局特征,分帧后再提取384维的语句级局部特征,并利用Fisher准则进行降维,最后提出一种弱尺度融合策略来将这两种语句级特征进行融合,再利用SVM进行情感分类。基于柏林情感库的实验结果表明本文方法较单纯使用语句级局部特征最后识别率提高了4.2%到13.8%,特别在小样本的情况下,语音情感识别率波动较小。
Speech is a short-term, stationary, time-frequency signal, so most researchers extract the emotional features by framing. However, the feature extracted after the frame segmentation is a local feature that can not accurately reflect the dynamic characteristics of the emotional speech, so that it is often unable to construct a robust emotion recognition system using only local features. In order to solve this problem, we first extract sentence-level global features by multi-scale optimal wavelet packet decomposition in frame-free speech signals, extract 384-dimensional sentence-level local features after frame segmentation, and use the Fisher criterion to reduce the dimension. Finally, A weak-scale fusion strategy to fuse these two sentence-level features, and then use SVM for emotion classification. Experimental results based on the Berlin Sentiment Library show that the proposed method improves the recognition rate by 4.2% to 13.8% compared with the simple use of sentence-level local features. Especially in the case of small samples, the speech affective recognition rate fluctuates less.