如何在scikit-learn GP核中固定超参数?自定义核梯度问题求助
Scikit-learn自定义GP核梯度返回问题解决
问题场景
我尝试在scikit-learn中实现以下简单的自定义GP核:
class uniform(Kernel): def __init__(self, M): # Initialize the parameters of your custom kernel self.M = M # Call the superclass constructor super().__init__() @property def hyperparameter_M(self): return Hyperparameter("M", "numeric", "fixed") def __call__(self, X, Y=None, eval_gradient=False): # The gradient of the kernel k(X, X) with respect to the log of the # hyperparameter of the kernel. Only returned when `eval_gradient` # is True. if Y is None: Y = X X = np.atleast_2d(X) Y = np.atleast_2d(Y) D = np.shape(X)[1] covariance_matrix = 1 w = np.arange(1,self.M+1,1) for d in range(D): a = feature(X[:,d],self.M).eval() b = feature(Y[:,d],self.M).eval() covariance_matrix_d = np.inner(a, b) covariance_matrix *= covariance_matrix_d if eval_gradient: # grad_matrix = np.zeros(covariance_matrix.shape) return covariance_matrix, None else: return covariance_matrix def diag(self, X): return np.diag(self.__call__(X, Y=X, eval_gradient=False)) def is_stationary(self): return True
使用该核拟合GP时,明明没有非固定超参数(uniform.theta是空列表),却仍被要求提供梯度,返回None也解决不了问题,哪里错了?
问题原因
scikit-learn的Kernel接口有严格规范:哪怕没有可优化的超参数,当eval_gradient=True时,也必须返回一个空的梯度列表,而不是None。你的M是固定超参数,属于不可训练的,但接口依然要求返回符合格式的结果,返回None会触发格式校验错误。
修正方案
1. 修正梯度返回值
把eval_gradient=True分支的返回值从None改成空列表[],这样就符合接口要求了。
2. 优化diag方法(可选)
原diag方法先计算完整协方差矩阵再取对角线,效率很低。从你的核逻辑来看,X=Y时对角线元素都是1,直接返回全1数组更高效。
修正后的完整代码:
class uniform(Kernel): def __init__(self, M): self.M = M super().__init__() @property def hyperparameter_M(self): return Hyperparameter("M", "numeric", "fixed") def __call__(self, X, Y=None, eval_gradient=False): if Y is None: Y = X X = np.atleast_2d(X) Y = np.atleast_2d(Y) D = X.shape[1] covariance_matrix = 1 for d in range(D): a = feature(X[:,d], self.M).eval() b = feature(Y[:,d], self.M).eval() covariance_matrix_d = np.inner(a, b) covariance_matrix *= covariance_matrix_d if eval_gradient: # 无优化超参数,返回空梯度列表 return covariance_matrix, [] else: return covariance_matrix def diag(self, X): # 直接返回全1数组,替代低效的取对角线操作 return np.ones(X.shape[0]) def is_stationary(self): return True
关键提示
- 只要继承了scikit-learn的
Kernel类,不管有没有可训练超参数,eval_gradient=True时都要返回(协方差矩阵, 梯度列表)的格式,梯度列表长度要和可训练超参数数量一致(这里是0,所以返回空列表)。 diag方法的优化不是必须的,但能显著提升小批量数据的处理速度。
内容的提问来源于stack exchange,提问作者Marco Rossi
相关产品推荐
相关产品推荐

