论文部分内容阅读
针对分布式存储系统的数据可用性问题展开了深入的研究,提出了一种支持纠删码的冗余倍数估计算法,根据数据统计特征获取单个数据块最优冗余方案;并基于该算法模型设计了一种适用于分布式存储系统的数据冗余策略,旨在消耗最小的存储开销获得最优的数据可用性.在实现该数据冗余策略的过程中,为了优化理论算法模型的工程可行性,提出了基于采样计算中间经验参数的方法,有效地利用目标存储数据的统计特征降低算法的计算复杂度.仿真实验验证了这种数据冗余策略的可行性和有效性.
Aiming at the problem of data availability of distributed storage system, an in-depth study has been carried out. A redundancy multiple estimation algorithm supporting erasure codes is proposed, and the optimal redundancy scheme of single data block is obtained according to the statistical characteristics of the data. Based on this algorithm model design A data redundancy strategy that is suitable for distributed storage system aims to consume the least storage overhead to obtain the optimal data availability.In the process of implementing this data redundancy strategy, in order to optimize the engineering feasibility of the theoretical algorithm model, A method based on sampling and calculating intermediate empirical parameters is proposed to effectively reduce the computational complexity of the algorithm by using the statistical characteristics of the target storage data. Simulation results verify the feasibility and effectiveness of this data redundancy strategy.