论文部分内容阅读
煤矿企业日志数据分析系统中,对数据的聚合统计,多数据源的数据关联是各种分析的基础操作。设计基于Hadoop的日志数据分析架构,并针对大数据关联准确率低的现象,将Bloom Filter算法与Map-Reduce处理架构融合。相比于传统的数据分析处理方式,算法具有容错性高、效率高等优点。
In the log data analysis system of coal mine enterprises, the data aggregation and statistics and the data association of multiple data sources are the basic operations of various analyzes. Design Hadoop-based log data analysis architecture, and Bloom Filter algorithm and Map-Reduce processing architecture for the phenomenon of low accuracy of large data association. Compared with the traditional data analysis and processing methods, the algorithm has the advantages of high fault tolerance and high efficiency.