You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras自定义层被拆解为多操作,无法获取权重的问题

自定义层权重获取问题及解决方案

问题背景

我想获取自定义层的权重,但无法通过model.layers[X].get_weights()实现。检查模型结构后发现,自定义层被拆解成了多个独立操作,这些子层里找不到权重。我希望把前5层替换成一个包含可训练kernel的单个层,能直接用get_weights()获取权重。

自定义层代码

class PixelBaseConv(Layer):

    def __init__(self, output_dim, **kwargs):
        self.output_dim = output_dim
        super(PixelBaseConv, self).__init__(**kwargs)

    def build(self, input_shape):
        # kernel_shape: w*h*c*output_dim
        kernel_size = input_shape[1:]
        kernel_shape = (1,) + kernel_size + (self.output_dim, )
        self.kernel = self.add_weight(name='kernel', 
                                      shape=kernel_shape,
                                      initializer='uniform',
                                      trainable=True)
        super(PixelBaseConv, self).build(input_shape)

    def call(self, inputs):
        # output_shape: w*h*output_dim
        outputs = []
        inputs = K.cast(inputs, dtype="float32")
        for i in range(self.output_dim):
            #output = tf.keras.layers.Multiply()([inputs, self.kernel[..., i]])
            output = inputs*self.kernel[...,i]
            output = K.sum(output, axis=-1)
            if len(outputs) != 0:
                outputs = np.dstack([outputs, output])
            else:
                outputs = output[..., np.newaxis]
        return tf.convert_to_tensor(outputs)

    def compute_output_shape(self, input_shape):
        return input_shape + (self.output_dim, )

模型结构异常表现

模型结构中,原本的PixelBaseConv层被拆解为多个类似tf.math.multiply、tf.math.reduce_sum、tf.concat等基础操作层,这些子层均无训练权重,导致无法通过常规方式获取自定义层的kernel权重。

尝试的代码及报错

我用以下代码遍历前10层的权重列表长度,并尝试打印第1层的权重:

for i in range(len(model.layers)):
    print("layer " + str(i), len(model.layers[i].get_weights()))
print(model.layers[1].get_weights()[0])

执行结果显示前10层的权重列表长度均为0,打印第1层权重时触发索引越界错误(IndexError: list index out of range),因为该层没有权重参数。

解决方案

问题出在自定义层的call方法中使用了numpy操作(np.dstack、np.newaxis),这会导致Keras无法将整个逻辑识别为一个完整的自定义层,而是拆解成多个基础操作节点。要修复这个问题,需要把所有numpy操作替换为TensorFlow/Keras的张量操作,确保层的逻辑完全在计算图内执行:

修改后的自定义层代码

class PixelBaseConv(Layer):

    def __init__(self, output_dim, **kwargs):
        self.output_dim = output_dim
        super(PixelBaseConv, self).__init__(**kwargs)

    def build(self, input_shape):
        # kernel_shape: 1*w*h*c*output_dim
        kernel_size = input_shape[1:]
        kernel_shape = (1,) + kernel_size + (self.output_dim, )
        self.kernel = self.add_weight(name='kernel', 
                                      shape=kernel_shape,
                                      initializer='uniform',
                                      trainable=True)
        super(PixelBaseConv, self).build(input_shape)

    def call(self, inputs):
        inputs = tf.cast(inputs, dtype="float32")
        # 用张量操作替代循环和numpy操作
        # 扩展input维度:(batch, w, h, c) -> (batch, w, h, c, 1)
        inputs_expanded = tf.expand_dims(inputs, axis=-1)
        # 逐通道相乘后求和:(batch, w, h, c, output_dim) -> (batch, w, h, output_dim)
        outputs = tf.reduce_sum(inputs_expanded * self.kernel, axis=-2)
        return outputs

    def compute_output_shape(self, input_shape):
        return input_shape[:-1] + (self.output_dim, )

修复说明

  1. 移除了循环和np.dstack,改用张量广播机制实现批量计算,效率更高且符合Keras层的规范
  2. 用tf.expand_dims替代np.newaxis,用tf.reduce_sum替代K.sum(两者功能一致,但保持TensorFlow API统一)
  3. 修正了compute_output_shape的返回值:原输入形状是(batch, w, h, c),输出应为(batch, w, h, output_dim),原代码多保留了输入的c维度,会导致形状不匹配

使用验证

替换自定义层后重新构建模型,此时PixelBaseConv会被识别为一个独立层,直接通过model.layers[X].get_weights()即可获取到kernel权重:

# 假设自定义层是模型的第1层
weights = model.layers[1].get_weights()
print(len(weights))  # 输出1,对应自定义层的kernel
print(weights[0].shape)  # 输出(1, w, h, c, output_dim),与定义的kernel_shape一致

内容的提问来源于stack exchange,提问作者sj L

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 01:55:23