TensorFlow many-to-many RNN时间序列二分类:xentropy替换失败
解决多对多RNN使用交叉熵损失的问题
我明白你现在遇到的问题了——把MSE换成交叉熵后模型跑不起来,这主要是因为分类任务的输出层设置、标签格式和交叉熵函数的要求不匹配。咱们一步步来修正:
1. 调整输出层维度
原来的n_outputs=1是回归任务的设置,现在你做的是二分类任务,需要输出2个类别的logits(对应标签0和1),所以要把n_outputs改成2,同时OutputProjectionWrapper的output_size也要对应修改:
n_outputs = 2 # 二分类,输出两个类别的logits cell = tf.contrib.rnn.OutputProjectionWrapper( tf.contrib.rnn.BasicRNNCell(num_units=n_neurons, activation=tf.nn.relu), output_size=n_outputs) # 这里也要改成2
2. 修正标签格式
tf.nn.sparse_softmax_cross_entropy_with_logits有两个关键要求:
- 标签必须是从0开始的整数类型(你的标签现在是1和2,要改成0和1)
- 标签的形状要和logits的前两个维度匹配:如果logits是
[batch_size, n_steps, n_outputs],那么标签应该是[batch_size, n_steps](不能有最后那个维度为1的形状)
所以先修改标签生成的代码:
# 修正标签:把<0设为0,>0设为1(从0开始的整数) dfX2 = np.array([np.sin(2*np.pi*f * (i/fs)) for i in x]) dfX2[dfX2 < 0] = 0 dfX2[dfX2 > 0] = 1 dfX2 = dfX2.astype(np.int32) # 转换成整数类型
然后在喂数据的时候,去掉标签的最后一个维度:
y_batch= dfX2[(iteration+1):(iteration+batch_size+1)] y_batch= y_batch.reshape(-1, n_steps) # 形状是[batch_size//n_steps, n_steps],去掉最后一个1的维度
同时,placeholder的类型也要改成整数:
y = tf.placeholder(tf.int32, [None, n_steps]) # 不再是[None, n_steps, n_outputs]
3. 修正交叉熵损失的调用
现在logits的形状是[batch_size, n_steps, 2],labels是[batch_size, n_steps],可以直接用sparse_softmax_cross_entropy_with_logits,注意要指定axis=-1(表示logits的最后维度是类别):
xentropy = tf.nn.sparse_softmax_cross_entropy_with_logits( labels=y, logits=outputs, axis=-1) # 指定类别所在的轴 loss = tf.reduce_mean(xentropy)
完整修改后的代码
把这些修改整合到你的代码里,完整版本如下:
import matplotlib.pyplot as plt import tensorflow as tf import numpy as np n_steps = 20 n_inputs = 1 n_neurons = 100 n_outputs = 2 # 改成2,对应二分类 learning_rate = 0.001 # Create data fs = 1000 # sample rate f = 2 # the frequency of the signal x = np.arange(fs) # the points on the x axis for plotting # training features dfX = np.array([ np.sin(2*np.pi*f * (i/fs)) for i in x]) # 修正标签:从0开始的整数类型 dfX2 = np.array([ np.sin(2*np.pi*f * (i/fs)) for i in x]) dfX2[dfX2 < 0] = 0 dfX2[dfX2 > 0] = 1 dfX2 = dfX2.astype(np.int32) # RNN X = tf.placeholder(tf.float32, [None, n_steps, n_inputs]) y = tf.placeholder(tf.int32, [None, n_steps]) # 修改标签的placeholder形状和类型 cell = tf.contrib.rnn.OutputProjectionWrapper( tf.contrib.rnn.BasicRNNCell(num_units=n_neurons, activation=tf.nn.relu), output_size=n_outputs) # output_size改成2 outputs, states = tf.nn.dynamic_rnn(cell, X, dtype=tf.float32) # 修正交叉熵损失 xentropy = tf.nn.sparse_softmax_cross_entropy_with_logits(labels=y, logits=outputs, axis=-1) loss = tf.reduce_mean(xentropy) optimizer = tf.train.AdamOptimizer(learning_rate=learning_rate) training_op = optimizer.minimize(loss) init = tf.global_variables_initializer() n_iterations = 1000 batch_size = 200 n_epouchs=10 with tf.Session() as sess: init.run() for epouch in range(n_epouchs): for iteration in range(n_iterations-200): X_batch= dfX[iteration:iteration+batch_size] X_batch= X_batch.reshape(-1,n_steps,n_inputs) y_batch= dfX2[(iteration+1):(iteration+batch_size+1)] y_batch= y_batch.reshape(-1,n_steps) # 去掉最后一个维度 sess.run(training_op, feed_dict={X: X_batch, y: y_batch}) if iteration % 100 == 0: current_loss = loss.eval(feed_dict={X: X_batch, y: y_batch}) print(iteration,"--",epouch, "\tLoss:", current_loss) # 预测示例:取logits中概率最大的类别 X_new1= dfX[37:37+batch_size] X_new1= X_new1.reshape(-1,n_steps,n_inputs) y_pred_logits = sess.run(outputs, feed_dict={X: X_new1}) y_pred = np.argmax(y_pred_logits, axis=-1) print("预测结果示例:", y_pred[0])
为什么原来的代码不行?
- 输出层维度不对:回归用1个输出,分类需要对应类别数的输出
- 标签格式不符合要求:sparse交叉熵要求标签是0开始的整数,且形状和logits的前N-1维匹配
- 损失函数参数不匹配:原来的y是float32且带最后一个维度,和logits的维度不兼容
这样修改后,你的多对多RNN就能用交叉熵损失来训练二分类任务啦!
内容的提问来源于stack exchange,提问作者aspire57
相关产品推荐
相关产品推荐

