TensorFlow 2.4.1嵌套自定义Layer导致可训练参数为0的原因咨询
call()? Let's break down exactly why your original basic_module was showing 0 trainable parameters, and why the fixed version works as expected.
The Core Issue: How Keras Tracks Trainable Parameters
Keras (powered by TensorFlow) keeps track of a Layer's trainable parameters by registering sub-layers that are initialized as instance attributes during the Layer's __init__ phase. When you build a custom Layer, any sub-Layer you assign to self.something in __init__ gets automatically added to the parent Layer's internal registry of child layers. This registry is what Keras uses to tally up trainable parameters for the model summary, and to ensure those parameters are included in training updates.
What Went Wrong in Your Original Code
In your initial basic_module:
class basic_module(Layer): def __init__(self, filters, kernel_size, strides): super(basic_module, self).__init__() self.res = basic_residual # You only saved the class reference, not an instance self.args = (filters, kernel_size, strides) def call(self, x): for _ in range(4): x = self.res(*self.args)(x) # Creates a NEW basic_residual instance EVERY time call() runs return x
Here's the breakdown of the problem:
- You never actually instantiated
basic_residualin__init__—you just stored the class itself, not a usable layer instance. - Every time
call()runs (during model building, inference, and training), you create a brand newbasic_residualinstance. These dynamic instances are never registered as child layers ofbasic_module—Keras has no way of knowing they exist, so their parameters don't show up in the model summary. - Beyond the parameter count issue, this approach is inefficient (creating new layers on every forward pass) and would break training entirely—each new layer has fresh, untrained parameters that never get updated across iterations.
Why the Fixed Version Works
Your modified code fixes this by moving all sub-layer initialization to __init__:
class basic_module(Layer): def __init__(self, filters, kernel_size, strides): super(basic_module, self).__init__() self.clayers = [basic_residual(filters, kernel_size, strides) for _ in range(4)] def call(self, x): for idx in range(4): x = self.clayers[idx](x) # Reuses the pre-initialized layers return x
- Now, you create all 4
basic_residualinstances once duringbasic_moduleinitialization, and store them as an instance attribute (self.clayers). - Keras automatically detects these sub-layers during the initialization phase, adds them to the parent Layer's registry, and includes their trainable parameters in the model summary.
- During
call(), you're reusing the same layer instances every time, so their parameters are consistent and get updated correctly during training.
A Quick Side Note
If you ever did need to create layers dynamically in call() (a rare edge case), you could manually register them with self.add_layer(new_layer) inside call(). But this is not recommended for your scenario—initializing sub-layers in __init__ is the standard, efficient, and reliable approach for reusable custom layers.
内容的提问来源于stack exchange,提问作者Will.Evo

