论文部分内容阅读
考虑汉语连续语音中的协同发音现象对语音识别性能的提高是非常重要的。针对汉语语音的特点,提出了一种新的在汉语连续语音识别中考虑音节间协同发音现象,对声学模型进行细化的识别单元。然后基于语音学知识对音节间上下文影响进行分类,实现单元间状态参数的共享,降低了模型的复杂程度,保证了模型的可训练度。这种方法和传统方法的最大不同在于:这种方法完全利用语音学知识进行聚类,而传统方法采用数据驱动的聚类方式。识别实验表明,基于语音学分类的音节间相关识别单元对识别性能有明显的改善,系统的首选误识率降低了17%。
It is very important to consider the improvement of speech recognition performance when considering the co-pronouncing phenomenon in Chinese continuous speech. Aiming at the characteristics of Chinese speech, this paper proposes a new recognition unit which takes into account the synergistic pronunciation between syllables and the acoustic model in Chinese continuous speech recognition. Then based on the phonetic knowledge, the context influence between syllables is classified to share the state parameters between the cells, which reduces the complexity of the model and ensures the training of the model. The biggest difference between this method and the traditional method is that this method completely uses phonetics knowledge for clustering, while the traditional method uses data-driven clustering. The recognition experiments show that the recognition units based on phonetic classification have obvious improvement on recognition performance, and the system’s preferred misclassification rate is reduced by 17%.