首次编写TensorFlow机器学习代码遇ValueError,求问题排查与修复
问题描述
首次尝试基于TensorFlow编写机器学习代码,构建了包含Embedding、双向LSTM、Bahdanau注意力等组件的文本模型,但调用model.fit()训练时触发ValueError,不清楚问题原因及修复方法。
模型代码
sequence_input = Input(shape=(np.array(padded_smishing).shape[1], ), dtype='float32') embedded_sequences = Embedding(Vocab_size, max_len, input_length=max_len, mask_zero= True)(sequence_input) lstm = Bidirectional(LSTM(64, dropout=0.5, return_sequences = True))(embedded_sequences) lstm, forward_h, forward_c, backward_h, backward_c = Bidirectional \ (LSTM(64, dropout=0.5, return_sequences=True, return_state=True))(lstm) state_h = Concatenate()([forward_h, backward_h]) # 隐藏状态 state_c = Concatenate()([forward_c, backward_c]) # 细胞状态 attention = BahdanauAttention(64) # 定义权重大小 context_vector, attention_weights = attention(lstm, state_h) z_s = WeightedSum()([embedded_sequences, attention_weights]) p_t = Dense(128)(z_s) p_t = Activation('softmax')(p_t) r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(tf.convert_to_tensor(p_t)) model = Model(inputs=sequence_input, outputs = r_s) mypotim = Adam(learning_rate = 0.001, beta_1 =0.9, beta_2 = 0.999, epsilon=1e-08, decay=0.0) model.compile(optimizer=mypotim,loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True), metrics=['accuracy'])
报错信息
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-170-d2314a49eedb> in <module> ----> 1 model.fit(X_train, y_train, batch_size=24, epochs=10, verbose=1, validation_data=(X_valid, y_valid), callbacks=callbacks) ValueError: tf.function only supports singleton tf.Variables created on the first call. Make sure the tf.Variable is only created once or created outside tf.function. See https://www.tensorflow.org/guide/function#creating_tfvariables for more information.
问题原因与修复方案
核心原因
报错根源是模型中的自定义层变量创建逻辑违反了TensorFlow函数式API的规则,具体有两点:
- 手动调用
tf.convert_to_tensor(p_t)传入自定义层WeightedAspectEmb,破坏了Keras层的变量追踪机制,导致变量重复创建。 - 如果
WeightedAspectEmb内部在call方法中创建可训练变量,而非在__init__或build方法中初始化,会触发tf.function的变量创建限制。
修复步骤
步骤1:移除手动张量转换
Keras层的输入本身就是张量,不需要手动转换,直接传入p_t即可:
# 原代码 r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(tf.convert_to_tensor(p_t)) # 修改后 r_s =WeightedAspectEmb(128, 128, W_constraint=MaxNorm(10), W_regularizer=ortho_reg)(p_t)
步骤2:修正自定义层WeightedAspectEmb的变量创建逻辑
确保可训练变量在build方法中初始化,而非call方法。示例实现如下:
class WeightedAspectEmb(tf.keras.layers.Layer): def __init__(self, input_dim, output_dim, W_constraint=None, W_regularizer=None, **kwargs): super().__init__(**kwargs) self.input_dim = input_dim self.output_dim = output_dim self.W_constraint = W_constraint self.W_regularizer = W_regularizer def build(self, input_shape): # 仅在第一次构建时创建可训练变量 self.W = self.add_weight( shape=(self.input_dim, self.output_dim), initializer='glorot_uniform', constraint=self.W_constraint, regularizer=self.W_regularizer, trainable=True, name='aspect_emb_weight' ) super().build(input_shape) def call(self, inputs): # call方法仅执行计算逻辑,不创建变量 return tf.matmul(inputs, self.W)
步骤3:匹配损失函数与模型输出特性
当前模型中p_t经过了softmax激活,而损失函数设置了from_logits=True(要求输入为未经过激活的logits),二者不匹配,需二选一修改:
- 方案一:修改损失函数参数
model.compile(optimizer=mypotim,loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=False), metrics=['accuracy'])
- 方案二:移除
p_t的softmax激活,让r_s输出logits
# 原代码 p_t = Activation('softmax')(p_t) # 修改后移除该行
内容的提问来源于stack exchange,提问作者skwldnjs1
相关产品推荐
相关产品推荐

