论文部分内容阅读
选取了258个苯酚类化合物的生物毒性数据,通过软件ADMEWORKS Model Builder的计算,选出7个结构描述符作为样本的结构参数,用稳健诊断方法剔除24个奇异样本,分别采用K最近邻方法和K均值聚类方法对剩余的234个样本数据进行分类,对分好的每一个类分别随机选择外部测试集,并用球型排除算法划分训练集和内部测试集,然后运用多元线性回归(Multiple Linear Regression,MLR)、偏最小二乘(Partial Least Squares,PLS)和人工神经网络(Artificial Neural Networks,ANN)方法进行预测模型的建立,计算结果表明,非线性模型的预测结果优于线性模型,有管理的分类方法(K nearest neighbors method,KNN)的预测结果优于无管理的分类方法(K均值聚类法)。
The biological toxicity data of 258 phenolic compounds were selected. Seven structural descriptors were selected as the structural parameters of the sample by the software ADMEWORKS Model Builder, and twenty-four singular samples were removed by using the robust diagnostic method. K nearest neighbor methods K-means clustering method is used to classify the remaining 234 sample data, and each outer class is randomly selected for each well. The training set and the internal test set are partitioned by the ball-type elimination algorithm. Then multiple linear regression Regression, MLR, Partial Least Squares (PLS) and Artificial Neural Networks (ANN) were used to predict the model. The results show that the prediction results of the nonlinear model are better than the linear model The K nearest neighbors method (KNN) predicts better than the unmanaged classification (K-means clustering).