基于TensorFlow的LSTM序列类别区间检测训练问题求助
Sequence Segment Classification: Troubleshooting & Resolution
Let's break down your sequence segment classification problem, the hurdles you hit, and the solution you landed on:
Problem Overview
You need to classify variable-length sequence segments (Class 1 typically covers timesteps ~14-16, Class 2 ~10-14) using an RNN-based framework with 6 distinct tasks:
- Binary classification for Class 1 presence (
[batch_size,2]) - Binary classification for Class 2 presence (
[batch_size,2]) - Class 1 start timestep prediction (
[batch_size,time_steps]) - Class 1 end timestep prediction (
[batch_size,time_steps]) - Class 2 start timestep prediction (
[batch_size,time_steps]) - Class 2 end timestep prediction (
[batch_size,time_steps])
Key Issues Encountered
- The first two binary classification tasks showed no meaningful improvement and remained stuck at random-guessing performance.
- The last four timestep prediction tasks had severe bias towards the 0 class.
Your Implementation Details
Data Generation Code
This function generates synthetic sequences with injected segments for Class 1 and 2:
def make_data(y = None,batch_size = 1, max_length = 100): f_name_values = [] l_name_values = [] X = np.random.randn(max_length,batch_size,1) if y is None: y_ = np.random.randint(0,2,(batch_size,2)) else: y_ = y y_1 = np.zeros((batch_size,4),dtype =int) y = np.concatenate([y_,y_1],axis = 1) seq_len = np.random.randint(int(max_length*0.2),int(max_length*0.6),(batch_size,1)) for i in range(batch_size): f_name = np.random.randint(0,seq_len[i,0]-3) if y[i,0] == 1: y[i,2], y[i,3] = int(f_name), int(f_name + np.random.randint(1,3)) X[y[i,2]:y[i,3],i,:] += 10 l_name_list = [j for j in range(seq_len[i,0]) if not( j <= y[i,3] and j > y[i,2]) ] l_name = np.random.choice(l_name_list) if y[i,1] == 1: y[i,4], y[i,5] = int(l_name), int(l_name + np.random.randint(1,4)) X[y[i,4]:y[i,5],i,:] += 10.5 return X, y , seq_len
Network Architecture
Multi-Layer LSTM Encoder
# Build RNN cell multi_encoder_cell = tf.contrib.rnn.MultiRNNCell([tf.contrib.rnn.BasicLSTMCell(num_units) for _ in range(layers)]) # Run Dynamic RNN # encoder_outputs: [max_time, batch_size, num_units] # encoder_state: [batch_size, num_units] # input: [max_time, batch_size, depth] w/ time_major= True encoder_outputs, encoder_state = tf.nn.dynamic_rnn( cell = multi_encoder_cell, inputs = X, sequence_length=s_sequence_length[:,0], time_major=True,dtype = 'float32')
Logit & Loss Calculation Function
This function maps RNN outputs to logits and computes task-specific loss (with masking for non-present classes):
def rnn_out_2_logit(scope_name, encoder_outputs_, y, num_of_classes,w,b,is_from_certain_class = None): with tf.variable_scope(scope_name + '_2_logit'): output_ = encoder_outputs_[-1,:,:] logit = tf.matmul(output_,w) + b if num_of_classes > 10: # Mask loss if the target class isn't present in the sequence bool_if_greater = tf.greater(is_from_certain_class,tf.zeros_like(is_from_certain_class)) int_bool = tf.cast(bool_if_greater,tf.float32) cross_entropy = tf.nn.sparse_softmax_cross_entropy_with_logits(labels = y, logits = logit) cross_entropy = cross_entropy * int_bool t_cross_entropy = tf.where(tf.less(tf.reduce_sum(int_bool), 1e-7),1e-7,(tf.reduce_mean(cross_entropy)/tf.reduce_sum(int_bool) )* batch_size) t_cross_entropy = tf.Print(t_cross_entropy, [t_cross_entropy, cross_entropy, int_bool],"cross_entrpy: ") else: t_cross_entropy = tf.reduce_mean( tf.nn.sparse_softmax_cross_entropy_with_logits(labels = y, logits = logit))
Task Function Calls
# Class 1 presence classification cross_entropy_1, logit_l, l, acc_l = rnn_out_2_logit("class_1",encoder_outputs,y[:,0], num_class_f_l, w_l, b_l) # Class 2 presence classification cross_entropy_2, logit_f, f, acc_f = rnn_out_2_logit("class_2",encoder_outputs,y[:,1], num_class_f_l, w_f, b_f) # Class 1 start timestep cross_entropy_3, logit_f_num_start, _f_num_start, acc_f_s = rnn_out_2_logit("class_1_start", encoder_outputs, y[:,2], max_words+1, w_l_e, b_l_e, y[:,1]) # Class 1 end timestep cross_entropy_4, logit_f_num_end, _f_num_end, acc_f_e = rnn_out_2_logit("class_1_end", encoder_outputs,y[:,3], max_words+1, w_f_e, b_f_e, y[:,1]) # Class 2 start timestep cross_entropy_5, logit_l_num_start, _l_num_start, acc_l_s = rnn_out_2_logit("class_2_start", encoder_outputs,y[:,4], max_words+1, w_l_s, b_l_s, y[:,0]) # Class 2 end timestep cross_entropy_6, logit_l_num_end, _l_num_end, acc_l_e = rnn_out_2_logit("class_2_end", encoder_outputs,y[:,5], max_words+1, w_l_e, b_l_e, y[:,0])
Training Observations
Sample Cross-Entropy Output (batch_size=3)
cross_entropy: [4.7751608][0 5.9443221 3.605999][0 1 1]
Training Logs (No Improvement)
epoch = 0 predicted y [[ 2.16714668 -1.0374999 ]] [[-0.30407697 0.01561086]] 99 3 35 99 true y [[ 1 1 21 22 13 15]] epoch = 100 predicted y [[ 1.26584482 -0.13619883]] [[-0.17700087 -0.11146521]] 99 3 35 99 true y [[ 1 1 5 6 32 33]] epoch = 200 predicted y [[ 0.70634729 0.42329881]] [[-0.1597006 -0.12876543]] 0 0 0 0 true y [[ 1 1 22 24 38 40]]
Progress & Resolution
- Update 1: Initially planned to split tasks: first resolve the binary classification problems, then switch to regression for timestep prediction.
- Update 2: Finally implemented a bidirectional RNN with time-step outputs, which delivered acceptable performance for all tasks.
内容的提问来源于stack exchange,提问作者NeoTT
相关产品推荐
相关产品推荐

