You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在scikit-learn GP核中固定超参数?自定义核梯度问题求助

Scikit-learn自定义GP核梯度返回问题解决

问题场景

我尝试在scikit-learn中实现以下简单的自定义GP核:

class uniform(Kernel):
    def __init__(self, M):
        # Initialize the parameters of your custom kernel
        self.M = M
        # Call the superclass constructor
        super().__init__()
        
    @property
    def hyperparameter_M(self):
        return Hyperparameter("M", "numeric", "fixed")
    
    def __call__(self, X, Y=None, eval_gradient=False):
        # The gradient of the kernel k(X, X) with respect to the log of the
        # hyperparameter of the kernel. Only returned when `eval_gradient`
        # is True.

        if Y is None:
            Y = X
        X = np.atleast_2d(X)
        Y = np.atleast_2d(Y)
        D = np.shape(X)[1]
        covariance_matrix = 1
        w = np.arange(1,self.M+1,1)
        for d in range(D):
            a = feature(X[:,d],self.M).eval()
            b = feature(Y[:,d],self.M).eval()
            covariance_matrix_d = np.inner(a, b)
            covariance_matrix *= covariance_matrix_d

        if eval_gradient:
            # grad_matrix = np.zeros(covariance_matrix.shape)
            return covariance_matrix, None
        else:
            return covariance_matrix

    def diag(self, X):
        return np.diag(self.__call__(X, Y=X, eval_gradient=False))
    
    def is_stationary(self):
        return True

使用该核拟合GP时,明明没有非固定超参数(uniform.theta是空列表),却仍被要求提供梯度,返回None也解决不了问题,哪里错了?

问题原因

scikit-learn的Kernel接口有严格规范:哪怕没有可优化的超参数,当eval_gradient=True时,也必须返回一个空的梯度列表,而不是None。你的M是固定超参数,属于不可训练的,但接口依然要求返回符合格式的结果,返回None会触发格式校验错误。

修正方案

1. 修正梯度返回值

把eval_gradient=True分支的返回值从None改成空列表[],这样就符合接口要求了。

2. 优化diag方法(可选)

原diag方法先计算完整协方差矩阵再取对角线,效率很低。从你的核逻辑来看,X=Y时对角线元素都是1,直接返回全1数组更高效。

修正后的完整代码:

class uniform(Kernel):
    def __init__(self, M):
        self.M = M
        super().__init__()
        
    @property
    def hyperparameter_M(self):
        return Hyperparameter("M", "numeric", "fixed")
    
    def __call__(self, X, Y=None, eval_gradient=False):
        if Y is None:
            Y = X
        X = np.atleast_2d(X)
        Y = np.atleast_2d(Y)
        D = X.shape[1]
        covariance_matrix = 1
        for d in range(D):
            a = feature(X[:,d], self.M).eval()
            b = feature(Y[:,d], self.M).eval()
            covariance_matrix_d = np.inner(a, b)
            covariance_matrix *= covariance_matrix_d

        if eval_gradient:
            # 无优化超参数,返回空梯度列表
            return covariance_matrix, []
        else:
            return covariance_matrix

    def diag(self, X):
        # 直接返回全1数组,替代低效的取对角线操作
        return np.ones(X.shape[0])
    
    def is_stationary(self):
        return True

关键提示

  • 只要继承了scikit-learn的Kernel类,不管有没有可训练超参数,eval_gradient=True时都要返回(协方差矩阵, 梯度列表)的格式,梯度列表长度要和可训练超参数数量一致(这里是0,所以返回空列表)。
  • diag方法的优化不是必须的,但能显著提升小批量数据的处理速度。

内容的提问来源于stack exchange,提问作者Marco Rossi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 13:54:55