论文部分内容阅读
[目的 /意义]旨在解决中文名称规范联合数据库检索系统CNASS的检索结果集记录量大且杂散的问题,实现其检索服务的关联聚簇功能。[方法 /过程]基于FRBR-LRM框架将个人名称规范记录转换为实体-属性-关系的RDF表示,利用记录内嵌的外部LC记录号重定向到VIAF记录,对原记录的作品关系等属性进行扩展。设计中文同名个人规范记录识别与聚簇算法,充分利用扩展后的作品关系,提高记录识别和聚簇的效率。[结果 /结论]选取300个人名,在CNASS中进行检索,对检索结果集运行算法,统计分析每个检索结果集的聚簇数和最大聚簇内记录数,综合计算聚簇效率指标,验证了本文聚簇算法的有效性。
[Purpose / Significance] aims to solve the problem of large recording volume and spurious search result set of CNASS in Chinese name specification, and to realize the related clustering function of its retrieval service. [Method / Procedure] Based on the FRBR-LRM framework, the personal name specification records are converted into entity-attribute-relational RDF representations, redirected to the VIAF records by using the external LC record numbers embedded in the records, and the attributes of the original record works Expand. The Chinese name recognition and clustering algorithm for personal records with the same name are designed to make full use of the expanded works relations to improve the efficiency of record recognition and clustering. [Result / Conclusion] The 300 names were selected and searched in CNASS. The algorithm was run on the retrieval result set, and the number of clusters in each retrieval result set and the maximum number of records in the cluster were statistically analyzed. The clustering efficiency index was calculated and verified The effectiveness of the clustering algorithm in this paper.