为何scipy.nnls与sklearn.LinearRegression(positive=True)结果不同?Super Learner场景
我在Python中实现自定义版本的Super Learner,使用三种方法计算基模型的权重:
- 自定义的
scipy.minimize(SLSQP方法,带权重和为1、0-1边界约束) scipy.nnls- 设置
positive=True的sklearn.LinearRegression
前两者结果相近,但LinearRegression的结果完全不同。查看源码发现当positive=True时,LinearRegression应该调用scipy.nnls,请问差异的原因是什么?
代码实现如下:
from sklearn.base import BaseEstimator, RegressorMixin from sklearn.utils.validation import check_X_y, check_array, check_is_fitted from sklearn.utils.estimator_checks import check_estimator from sklearn.model_selection import KFold from sklearn.pipeline import make_pipeline from sklearn.preprocessing import StandardScaler from sklearn import linear_model from sklearn import neighbors from sklearn import datasets import matplotlib.pyplot as plt from scipy import optimize from pandas.plotting import scatter_matrix import numpy as np import pandas as pd class SuperLearner(BaseEstimator, RegressorMixin): def __init__(self, base_estimators): self.base_estimators = base_estimators self.meta_learner = linear_model.LinearRegression(positive=True) self.weights = None def rss(self, weights, X, y): y_pred = np.dot(X, weights) return np.sum((y - y_pred)**2) def constraint(self, weights): return np.sum(weights) - 1 def fit(self, X, y): X, y = check_X_y(X, y) meta_predictions = np.zeros((X.shape[0], len(self.base_estimators)), dtype=np.float64) #TODO: modify the number of folds depending on the number of base estimators and the size of the dataset kf = KFold(n_splits=5) for i, (tran_idx, val_idx) in enumerate(kf.split(X)): X_train, X_val = X[tran_idx], X[val_idx] y_train, y_val = y[tran_idx], y[val_idx] for j, estimator in enumerate(self.base_estimators): estimator.fit(X_train, y_train) meta_predictions[val_idx, j] = estimator.predict(X_val) guess = np.empty(len(self.base_estimators)) bounds = [(0,1)] * len(self.base_estimators) result = optimize.minimize(self.rss, guess, args=(meta_predictions, y), method='SLSQP', bounds=bounds, constraints={'type':'eq', 'fun':self.constraint}) print(result.x, np.sum(result.x)) result = optimize.nnls(meta_predictions, y) print(result[0], np.sum(result[0])) self.meta_learner.fit(meta_predictions, y) self.weights= self.meta_learner.coef_ self.weights= self.weights / np.sum(self.weights) print(self.weights, np.sum(self.weights)) return self def predict(self, X): check_is_fitted(self, 'meta_learner') X = check_array(X) base_predictions = np.zeros((X.shape[0], len(self.base_estimators)), dtype=np.float64) for i, estimator in enumerate(self.base_estimators): base_predictions[:, i] = estimator.predict(X) return np.dot(base_predictions, self.weights) def main(): np.random.seed(100) X, y = datasets.make_friedman1(1000) ols = linear_model.LinearRegression() elastic = linear_model.ElasticNetCV() ridge = linear_model.RidgeCV() lars = linear_model.LarsCV() lasso = linear_model.LassoCV() knn = neighbors.KNeighborsRegressor() superLeaner = SuperLearner([ols, elastic, ridge, lars, lasso, knn]) superLeaner.fit(X, y) y_pred = superLeaner.predict(X) print("MSE: ", np.mean((y_pred - y)**2)) if __name__ == "__main__": main()
原因分析
差异的核心在于你对LinearRegression的结果做了额外的归一化操作,而另外两种方法的输出状态不同:
自定义SLSQP的约束特性:你给SLSQP设置了
np.sum(weights) - 1 = 0的等式约束,所以优化得到的权重本身就满足和为1的条件,不需要额外处理。打印的result.x直接是归一化后的权重。scipy.nnls的原始输出:nnls只保证权重非负,不做和为1的约束。你直接打印了它的原始结果,这些权重的和不等于1,和SLSQP的结果不在同一尺度上,但如果对nnls的结果做同样的
/ np.sum(weights)归一化,会发现它和LinearRegression处理后的结果完全一致。LinearRegression的后处理:当
positive=True时,LinearRegression确实调用了nnls得到原始非负权重,但你紧接着执行了self.weights= self.weights / np.sum(self.weights),把权重归一化到和为1,这才导致它看起来和未处理的nnls结果差异大——本质上两者的原始权重是一样的,只是你对LinearRegression的结果做了归一化。
简单验证:把nnls的结果也做归一化,比如:
w_nnls = result[0] w_nnls_normalized = w_nnls / np.sum(w_nnls) print(w_nnls_normalized)
输出会和你处理后的self.weights完全相同。
内容的提问来源于stack exchange,提问作者Dragos Tanasa

