论文部分内容阅读
在计算机进行现代汉语复句书读前后非分句语言片段的自动识别过程中,我们发现并总结出一些可形式化后供其执行的句法规则。这些规则的效用如何,我们还没来得及进行试验,本文也暂时未做分析。我们的设想是:待计算机工作人员将它们形式化为可供计算机理解和执行的语言后,在训练集内进行小规模的试验,再进而把试验范围扩大到整个语料库,从中不断改进和完善规则。
In the process of computer automatic recognition of non-clause language fragments before and after the modern Chinese complex sentence reading, we find and summarize some syntactic rules that can be formalized for their implementation. The effectiveness of these rules, we have not had time to experiment, this article has not yet analyzed. Our idea is that after the computer staff formalize them into a language that computers can understand and implement, small-scale experiments are conducted in the training set, and then the experiment is extended to the entire corpus to continuously improve and perfect the rules .