You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TensorFlow的LSTM序列类别区间检测训练问题求助

Sequence Segment Classification: Troubleshooting & Resolution

Let's break down your sequence segment classification problem, the hurdles you hit, and the solution you landed on:

Problem Overview

You need to classify variable-length sequence segments (Class 1 typically covers timesteps ~14-16, Class 2 ~10-14) using an RNN-based framework with 6 distinct tasks:

  • Binary classification for Class 1 presence ([batch_size,2])
  • Binary classification for Class 2 presence ([batch_size,2])
  • Class 1 start timestep prediction ([batch_size,time_steps])
  • Class 1 end timestep prediction ([batch_size,time_steps])
  • Class 2 start timestep prediction ([batch_size,time_steps])
  • Class 2 end timestep prediction ([batch_size,time_steps])

Key Issues Encountered

  • The first two binary classification tasks showed no meaningful improvement and remained stuck at random-guessing performance.
  • The last four timestep prediction tasks had severe bias towards the 0 class.

Your Implementation Details

Data Generation Code

This function generates synthetic sequences with injected segments for Class 1 and 2:

def make_data(y = None,batch_size = 1, max_length = 100):
    f_name_values = []
    l_name_values = []
    X = np.random.randn(max_length,batch_size,1)
    if y is None:
        y_ = np.random.randint(0,2,(batch_size,2))
    else:
        y_ = y
    y_1 = np.zeros((batch_size,4),dtype =int)
    y = np.concatenate([y_,y_1],axis = 1)
    seq_len = np.random.randint(int(max_length*0.2),int(max_length*0.6),(batch_size,1))
    for i in range(batch_size):
        f_name = np.random.randint(0,seq_len[i,0]-3)
        if y[i,0] == 1:
            y[i,2], y[i,3] = int(f_name), int(f_name + np.random.randint(1,3))
            X[y[i,2]:y[i,3],i,:] += 10
        l_name_list = [j for j in range(seq_len[i,0]) if not( j <= y[i,3] and j > y[i,2]) ]
        l_name = np.random.choice(l_name_list)
        if y[i,1] == 1:
            y[i,4], y[i,5] = int(l_name), int(l_name + np.random.randint(1,4))
            X[y[i,4]:y[i,5],i,:] += 10.5
    return X, y , seq_len

Network Architecture

Multi-Layer LSTM Encoder

# Build RNN cell
multi_encoder_cell = tf.contrib.rnn.MultiRNNCell([tf.contrib.rnn.BasicLSTMCell(num_units) for _ in range(layers)])
# Run Dynamic RNN
# encoder_outputs: [max_time, batch_size, num_units]
# encoder_state: [batch_size, num_units]
# input: [max_time, batch_size, depth] w/ time_major= True
encoder_outputs, encoder_state = tf.nn.dynamic_rnn(
    cell = multi_encoder_cell,
    inputs = X,
    sequence_length=s_sequence_length[:,0],
    time_major=True,dtype = 'float32')

Logit & Loss Calculation Function

This function maps RNN outputs to logits and computes task-specific loss (with masking for non-present classes):

def rnn_out_2_logit(scope_name, encoder_outputs_, y, num_of_classes,w,b,is_from_certain_class = None):
    with tf.variable_scope(scope_name + '_2_logit'):
        output_ = encoder_outputs_[-1,:,:]
        logit = tf.matmul(output_,w) + b
        if num_of_classes > 10:
            # Mask loss if the target class isn't present in the sequence
            bool_if_greater = tf.greater(is_from_certain_class,tf.zeros_like(is_from_certain_class))
            int_bool = tf.cast(bool_if_greater,tf.float32)
            cross_entropy = tf.nn.sparse_softmax_cross_entropy_with_logits(labels = y, logits = logit)
            cross_entropy = cross_entropy * int_bool
            t_cross_entropy = tf.where(tf.less(tf.reduce_sum(int_bool), 1e-7),1e-7,(tf.reduce_mean(cross_entropy)/tf.reduce_sum(int_bool) )* batch_size)
            t_cross_entropy = tf.Print(t_cross_entropy, [t_cross_entropy, cross_entropy, int_bool],"cross_entrpy: ")
        else:
            t_cross_entropy = tf.reduce_mean( tf.nn.sparse_softmax_cross_entropy_with_logits(labels = y, logits = logit))

Task Function Calls

# Class 1 presence classification
cross_entropy_1, logit_l, l, acc_l = rnn_out_2_logit("class_1",encoder_outputs,y[:,0], num_class_f_l, w_l, b_l)
# Class 2 presence classification
cross_entropy_2, logit_f, f, acc_f = rnn_out_2_logit("class_2",encoder_outputs,y[:,1], num_class_f_l, w_f, b_f)
# Class 1 start timestep
cross_entropy_3, logit_f_num_start, _f_num_start, acc_f_s = rnn_out_2_logit("class_1_start", encoder_outputs, y[:,2], max_words+1, w_l_e, b_l_e, y[:,1])
# Class 1 end timestep
cross_entropy_4, logit_f_num_end, _f_num_end, acc_f_e = rnn_out_2_logit("class_1_end", encoder_outputs,y[:,3], max_words+1, w_f_e, b_f_e, y[:,1])
# Class 2 start timestep
cross_entropy_5, logit_l_num_start, _l_num_start, acc_l_s = rnn_out_2_logit("class_2_start", encoder_outputs,y[:,4], max_words+1, w_l_s, b_l_s, y[:,0])
# Class 2 end timestep
cross_entropy_6, logit_l_num_end, _l_num_end, acc_l_e = rnn_out_2_logit("class_2_end", encoder_outputs,y[:,5], max_words+1, w_l_e, b_l_e, y[:,0])

Training Observations

Sample Cross-Entropy Output (batch_size=3)

cross_entropy: [4.7751608][0 5.9443221 3.605999][0 1 1]

Training Logs (No Improvement)

epoch = 0
predicted y [[ 2.16714668 -1.0374999 ]] [[-0.30407697 0.01561086]] 99 3 35 99
true y [[ 1 1 21 22 13 15]]
epoch = 100
predicted y [[ 1.26584482 -0.13619883]] [[-0.17700087 -0.11146521]] 99 3 35 99
true y [[ 1 1 5 6 32 33]]
epoch = 200
predicted y [[ 0.70634729 0.42329881]] [[-0.1597006 -0.12876543]] 0 0 0 0
true y [[ 1 1 22 24 38 40]]

Progress & Resolution

  • Update 1: Initially planned to split tasks: first resolve the binary classification problems, then switch to regression for timestep prediction.
  • Update 2: Finally implemented a bidirectional RNN with time-step outputs, which delivered acceptable performance for all tasks.

内容的提问来源于stack exchange,提问作者NeoTT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:41:45