You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

首次编写TensorFlow机器学习代码遇ValueError,求问题排查与修复

问题描述

首次尝试基于TensorFlow编写机器学习代码,构建了包含Embedding、双向LSTM、Bahdanau注意力等组件的文本模型,但调用model.fit()训练时触发ValueError,不清楚问题原因及修复方法。

模型代码

sequence_input = Input(shape=(np.array(padded_smishing).shape[1], ), dtype='float32')
embedded_sequences = Embedding(Vocab_size, max_len, input_length=max_len, mask_zero= True)(sequence_input)

lstm = Bidirectional(LSTM(64, dropout=0.5, return_sequences = True))(embedded_sequences)
lstm, forward_h, forward_c, backward_h, backward_c = Bidirectional \
  (LSTM(64, dropout=0.5, return_sequences=True, return_state=True))(lstm)

state_h = Concatenate()([forward_h, backward_h]) # 隐藏状态
state_c = Concatenate()([forward_c, backward_c]) # 细胞状态
attention = BahdanauAttention(64) # 定义权重大小
context_vector, attention_weights = attention(lstm, state_h)

z_s = WeightedSum()([embedded_sequences, attention_weights])
p_t = Dense(128)(z_s)
p_t = Activation('softmax')(p_t)
r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(tf.convert_to_tensor(p_t))

model = Model(inputs=sequence_input, outputs = r_s)

mypotim = Adam(learning_rate = 0.001, beta_1 =0.9, beta_2 = 0.999, epsilon=1e-08, decay=0.0)
model.compile(optimizer=mypotim,loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True), metrics=['accuracy'])

报错信息

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-170-d2314a49eedb> in <module>
----> 1 model.fit(X_train, y_train, batch_size=24, epochs=10, verbose=1, validation_data=(X_valid, y_valid), callbacks=callbacks)

    ValueError: tf.function only supports singleton tf.Variables created on the first call. Make sure the tf.Variable is only created once or created outside tf.function. See https://www.tensorflow.org/guide/function#creating_tfvariables for more information.
问题原因与修复方案

核心原因

报错根源是模型中的自定义层变量创建逻辑违反了TensorFlow函数式API的规则,具体有两点:

  1. 手动调用tf.convert_to_tensor(p_t)传入自定义层WeightedAspectEmb,破坏了Keras层的变量追踪机制,导致变量重复创建。
  2. 如果WeightedAspectEmb内部在call方法中创建可训练变量,而非在__init__或build方法中初始化,会触发tf.function的变量创建限制。

修复步骤

步骤1:移除手动张量转换

Keras层的输入本身就是张量,不需要手动转换,直接传入p_t即可:

# 原代码
r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(tf.convert_to_tensor(p_t))
# 修改后
r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(p_t)

步骤2:修正自定义层WeightedAspectEmb的变量创建逻辑

确保可训练变量在build方法中初始化,而非call方法。示例实现如下:

class WeightedAspectEmb(tf.keras.layers.Layer):
    def __init__(self, input_dim, output_dim, W_constraint=None, W_regularizer=None, **kwargs):
        super().__init__(**kwargs)
        self.input_dim = input_dim
        self.output_dim = output_dim
        self.W_constraint = W_constraint
        self.W_regularizer = W_regularizer

    def build(self, input_shape):
        # 仅在第一次构建时创建可训练变量
        self.W = self.add_weight(
            shape=(self.input_dim, self.output_dim),
            initializer='glorot_uniform',
            constraint=self.W_constraint,
            regularizer=self.W_regularizer,
            trainable=True,
            name='aspect_emb_weight'
        )
        super().build(input_shape)

    def call(self, inputs):
        # call方法仅执行计算逻辑,不创建变量
        return tf.matmul(inputs, self.W)

步骤3:匹配损失函数与模型输出特性

当前模型中p_t经过了softmax激活,而损失函数设置了from_logits=True(要求输入为未经过激活的logits),二者不匹配,需二选一修改:

  • 方案一:修改损失函数参数
model.compile(optimizer=mypotim,loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=False), metrics=['accuracy'])
  • 方案二:移除p_t的softmax激活,让r_s输出logits
# 原代码
p_t = Activation('softmax')(p_t)
# 修改后移除该行

内容的提问来源于stack exchange,提问作者skwldnjs1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:40:34