TensorFlow dynamic_rnn中sequence_length参数使用异常问题排查
问题诊断与修复方案
我仔细看了你的代码和问题描述,发现两个核心错误导致使用sequence_length时模型无法正常学习:
1. sequence_length参数传递错误
你当前传递给len_ph的是[len(example)],但example的形状是(1, seq_len, 1),len(example)返回的是第一个维度的大小(也就是1),这意味着TensorFlow认为所有序列的长度都是1,只会处理每个序列的第一个时间步,完全忽略了后续的时间步信息。
正确的做法是获取序列的实际时间步长度,也就是example.shape[1],喂数据时应该写成:
vals={data_ph: example, len_ph: [example.shape[1]], y_ph: label, }
2. 损失函数的输入形状与逻辑错误
你的任务是序列分类(每个序列对应一个类别标签),但:
- 你把
y_ph定义成了[1, None, 3],这是序列标注任务的形状(每个时间步对应一个标签),而非序列分类的[1, 3]; tf.nn.softmax_cross_entropy_with_logits要求logits是未经过softmax激活的原始输出,你现在把经过softmax的prediction传进去了,这会导致损失计算逻辑完全错误。
修正方案:
- 重新定义
y_ph为序列分类的形状[1, 3]; - 将logits和prediction分开,用原始输出计算损失,用softmax后的结果做预测。
修正后的完整代码
from __future__ import print_function import tensorflow as tf import numpy as np import random dataset = [[1, 0], [2, 0], [1,2], [1,1]] labels = [[1,0,0], [0,1,0], [0,1,0], [0,0,1]] #--------------------------------------------- #define model # placeholders data_ph = tf.placeholder("float", [1, None, 1], name="data_placeholder") len_ph = tf.placeholder("int32", [1], name="seq_len_placeholder") # 修正:y_ph改为序列分类的形状[1, 3] y_ph = tf.placeholder("float", [1, 3], name="y_placeholder") n_hidden = 10 n_out = len(labels[0]) # variable definition out_weights=tf.Variable(tf.random_normal([n_hidden,n_out])) out_bias=tf.Variable(tf.random_normal([n_out])) # lstm definition lstm_cell = tf.nn.rnn_cell.BasicLSTMCell(n_hidden, state_is_tuple=True) state_series, final_state = tf.nn.dynamic_rnn( cell=lstm_cell, inputs=data_ph, dtype=tf.float32, sequence_length=len_ph, time_major=False ) out = state_series[:, -1, :] # 修正:分离logits和prediction,logits用于损失计算,prediction用于预测 logits = tf.matmul(out, out_weights)+out_bias prediction=tf.nn.softmax(logits) # 修正:用logits计算损失,且labels形状匹配 loss = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=logits,labels=y_ph)) optimizer=tf.train.AdamOptimizer(learning_rate=1e-3).minimize(loss) #--------------------------------------------- #run model sess = tf.InteractiveSession() sess.run(tf.global_variables_initializer()) #TRAIN for iteration in range(5000): if (iteration%100 == 0): print(iteration) ind = random.randint(0, len(dataset)-1) example = np.reshape(dataset[ind], (1,-1,1)) # 修正:label的形状改为[1,3] label = np.reshape(labels[ind], (1,3)) # 修正:传递正确的序列长度 vals={data_ph: example, len_ph: [example.shape[1]], y_ph: label, } sess.run(optimizer, feed_dict=vals) #TEST for x in range(len(dataset)): example = np.reshape(dataset[x], (1,-1,1)) label = np.reshape(labels[x], (1,3)) vals = {data_ph: example, len_ph: [example.shape[1]], y_ph: label, } classification = sess.run([prediction, loss], feed_dict=vals) print("predicted values: "+str(np.matrix.round(classification[0][0], decimals=2)), "loss: "+str(classification[1]))
修正后的预期结果
运行修正后的代码,你会得到和不使用sequence_length时类似的正确预测结果,模型能够学习到你定义的分类规则:
- 序列只有1 → 分类为1
- 序列存在2 → 分类为2
- 序列前两个时间步都是1 → 分类为3
内容的提问来源于stack exchange,提问作者The_Chicken_Lord
相关产品推荐
相关产品推荐

