You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 2.4.1嵌套自定义Layer导致可训练参数为0的原因咨询

Why does my custom Keras Layer show 0 trainable parameters when creating sub-layers in call()?

Let's break down exactly why your original basic_module was showing 0 trainable parameters, and why the fixed version works as expected.

The Core Issue: How Keras Tracks Trainable Parameters

Keras (powered by TensorFlow) keeps track of a Layer's trainable parameters by registering sub-layers that are initialized as instance attributes during the Layer's __init__ phase. When you build a custom Layer, any sub-Layer you assign to self.something in __init__ gets automatically added to the parent Layer's internal registry of child layers. This registry is what Keras uses to tally up trainable parameters for the model summary, and to ensure those parameters are included in training updates.

What Went Wrong in Your Original Code

In your initial basic_module:

class basic_module(Layer):
    def __init__(self, filters, kernel_size, strides):
        super(basic_module, self).__init__()
        self.res = basic_residual  # You only saved the class reference, not an instance
        self.args = (filters, kernel_size, strides)
    def call(self, x):
        for _ in range(4):
            x = self.res(*self.args)(x)  # Creates a NEW basic_residual instance EVERY time call() runs
        return x

Here's the breakdown of the problem:

  1. You never actually instantiated basic_residual in __init__—you just stored the class itself, not a usable layer instance.
  2. Every time call() runs (during model building, inference, and training), you create a brand new basic_residual instance. These dynamic instances are never registered as child layers of basic_module—Keras has no way of knowing they exist, so their parameters don't show up in the model summary.
  3. Beyond the parameter count issue, this approach is inefficient (creating new layers on every forward pass) and would break training entirely—each new layer has fresh, untrained parameters that never get updated across iterations.

Why the Fixed Version Works

Your modified code fixes this by moving all sub-layer initialization to __init__:

class basic_module(Layer):
    def __init__(self, filters, kernel_size, strides):
        super(basic_module, self).__init__()
        self.clayers = [basic_residual(filters, kernel_size, strides) for _ in range(4)]
    def call(self, x):
        for idx in range(4):
            x = self.clayers[idx](x)  # Reuses the pre-initialized layers
        return x
  • Now, you create all 4 basic_residual instances once during basic_module initialization, and store them as an instance attribute (self.clayers).
  • Keras automatically detects these sub-layers during the initialization phase, adds them to the parent Layer's registry, and includes their trainable parameters in the model summary.
  • During call(), you're reusing the same layer instances every time, so their parameters are consistent and get updated correctly during training.

A Quick Side Note

If you ever did need to create layers dynamically in call() (a rare edge case), you could manually register them with self.add_layer(new_layer) inside call(). But this is not recommended for your scenario—initializing sub-layers in __init__ is the standard, efficient, and reliable approach for reusable custom layers.

内容的提问来源于stack exchange,提问作者Will.Evo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 14:17:26