论文部分内容阅读
交通信息基础数据元与用户数据项的中文名称短语的对应是数据元建立、标准符合性检测等工作的基础。为了提高名称对应的准确率,提出了一种利用数据元名称组成的特定结构进行数据项名称与数据元名称进行对应的方法,并给出了相似度的计算算法。该算法将用户数据项名称短语的省略情况按照中文语言习惯进行总结,采用数学中干扰修正的思想,分别按照语素和词素对相似度值进行计算,并利用相同语素的个数对相似度进行修正,综合得出词语的相似度。最后利用交通运输部实际工程数据进行了验证。研究结果表明:本算法较文献[1]中算法的“有改善”率提升了91.20%,“明显改善”率提升了9.62%;较文献[2]中的“有改善”率提升了88.40%,“明显改善”率提升了66.80%。
The correspondence between the basic information element of traffic information and the Chinese name phrase of user data item is the basis of data element establishment and standard compliance check. In order to improve the accuracy of name correspondence, a method of using the specific structure composed of data element names to match the names of data items and data element names is proposed, and the algorithm of similarity calculation is given. The algorithm summarizes the omission of the name phrase of user data according to the Chinese language habit. Based on the idea of interference correction in mathematics, the similarity value is calculated respectively according to morpheme and morpheme, and the similarity is corrected by the number of the same morpheme , Get the similarity of words. Finally, the actual engineering data of Ministry of Transport was used to verify. The results show that the proposed algorithm has a 91.20% improvement over the improved algorithm in the literature [1] and a 9.62% improvement in the “significantly improved” ratio compared with the algorithm in [2] “Rate increased by 88.40%, ” significantly improved "rate increased by 66.80%.