Keras自定义层被拆解为多操作,无法获取权重的问题
自定义层权重获取问题及解决方案
问题背景
我想获取自定义层的权重,但无法通过model.layers[X].get_weights()实现。检查模型结构后发现,自定义层被拆解成了多个独立操作,这些子层里找不到权重。我希望把前5层替换成一个包含可训练kernel的单个层,能直接用get_weights()获取权重。
自定义层代码
class PixelBaseConv(Layer): def __init__(self, output_dim, **kwargs): self.output_dim = output_dim super(PixelBaseConv, self).__init__(**kwargs) def build(self, input_shape): # kernel_shape: w*h*c*output_dim kernel_size = input_shape[1:] kernel_shape = (1,) + kernel_size + (self.output_dim, ) self.kernel = self.add_weight(name='kernel', shape=kernel_shape, initializer='uniform', trainable=True) super(PixelBaseConv, self).build(input_shape) def call(self, inputs): # output_shape: w*h*output_dim outputs = [] inputs = K.cast(inputs, dtype="float32") for i in range(self.output_dim): #output = tf.keras.layers.Multiply()([inputs, self.kernel[..., i]]) output = inputs*self.kernel[...,i] output = K.sum(output, axis=-1) if len(outputs) != 0: outputs = np.dstack([outputs, output]) else: outputs = output[..., np.newaxis] return tf.convert_to_tensor(outputs) def compute_output_shape(self, input_shape): return input_shape + (self.output_dim, )
模型结构异常表现
模型结构中,原本的PixelBaseConv层被拆解为多个类似tf.math.multiply、tf.math.reduce_sum、tf.concat等基础操作层,这些子层均无训练权重,导致无法通过常规方式获取自定义层的kernel权重。
尝试的代码及报错
我用以下代码遍历前10层的权重列表长度,并尝试打印第1层的权重:
for i in range(len(model.layers)): print("layer " + str(i), len(model.layers[i].get_weights())) print(model.layers[1].get_weights()[0])
执行结果显示前10层的权重列表长度均为0,打印第1层权重时触发索引越界错误(IndexError: list index out of range),因为该层没有权重参数。
解决方案
问题出在自定义层的call方法中使用了numpy操作(np.dstack、np.newaxis),这会导致Keras无法将整个逻辑识别为一个完整的自定义层,而是拆解成多个基础操作节点。要修复这个问题,需要把所有numpy操作替换为TensorFlow/Keras的张量操作,确保层的逻辑完全在计算图内执行:
修改后的自定义层代码
class PixelBaseConv(Layer): def __init__(self, output_dim, **kwargs): self.output_dim = output_dim super(PixelBaseConv, self).__init__(**kwargs) def build(self, input_shape): # kernel_shape: 1*w*h*c*output_dim kernel_size = input_shape[1:] kernel_shape = (1,) + kernel_size + (self.output_dim, ) self.kernel = self.add_weight(name='kernel', shape=kernel_shape, initializer='uniform', trainable=True) super(PixelBaseConv, self).build(input_shape) def call(self, inputs): inputs = tf.cast(inputs, dtype="float32") # 用张量操作替代循环和numpy操作 # 扩展input维度:(batch, w, h, c) -> (batch, w, h, c, 1) inputs_expanded = tf.expand_dims(inputs, axis=-1) # 逐通道相乘后求和:(batch, w, h, c, output_dim) -> (batch, w, h, output_dim) outputs = tf.reduce_sum(inputs_expanded * self.kernel, axis=-2) return outputs def compute_output_shape(self, input_shape): return input_shape[:-1] + (self.output_dim, )
修复说明
- 移除了循环和
np.dstack,改用张量广播机制实现批量计算,效率更高且符合Keras层的规范 - 用
tf.expand_dims替代np.newaxis,用tf.reduce_sum替代K.sum(两者功能一致,但保持TensorFlow API统一) - 修正了
compute_output_shape的返回值:原输入形状是(batch, w, h, c),输出应为(batch, w, h, output_dim),原代码多保留了输入的c维度,会导致形状不匹配
使用验证
替换自定义层后重新构建模型,此时PixelBaseConv会被识别为一个独立层,直接通过model.layers[X].get_weights()即可获取到kernel权重:
# 假设自定义层是模型的第1层 weights = model.layers[1].get_weights() print(len(weights)) # 输出1,对应自定义层的kernel print(weights[0].shape) # 输出(1, w, h, c, output_dim),与定义的kernel_shape一致
内容的提问来源于stack exchange,提问作者sj L
相关产品推荐
相关产品推荐

