论文部分内容阅读
Fluid office documents,as semi-structured data often represented by XML,are important parts of Big Data.These office documents have different formats,and their matching APIs depend on developing platform and versions,causing difficulty in custom development and information retrieval from them.To solve this problem,we have been developing an office document query (ODQ) language which provides a uniform method to retrieve content from documents with different formats and versions.ODQ builds common document model ontology to conceal the format details of documents and provides a uniform operation interface to handle office documents with different formats.The results show that ODQ has advantages in format independence,and can facilitate users in developing documents processing systems with good interoperability.