You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于单标签训练集的深度学习多标签时间序列分类方案问询

问题

我正在构建一个用于一维单变量时间序列分类的深度学习模型。训练数据集包含4类数据,每类数据量均衡;另有独立的测试数据集,但测试数据存在多标签样本。

训练数据集示例

[3. 3. 6. 7. 8. 2. 1. 1. 1. 3.] [1 0 0 0]
[7. 6. 4. 4. 6. 8. 3. 4. 6. 4.] [0 0 1 0]
[5. 5. 1. 7. 6. 3. 6. 1. 4. 2.] [1 0 0 0]
[1. 1. 5. 5. 7. 3. 4. 6. 4. 4.] [0 0 0 1]
[4. 2. 3. 1. 7. 2. 4. 1. 1. 7.] [0 1 0 0]

测试数据集示例

[ 4.  3.  3.  2.  5.  2.  3.  7.  2.  0.] [1 0 1 0]
[ 3.  3.  3.  1.  6.  3.  4.  4.  1.  3.] [0 1 0 0]
[ 2.  1.  7.  8.  4.  5.  0.  4.  6.  6.] [1 1 0 0]
[ 5.  1.  9.  5.  3.  4.  1.  8.  1.  8.] [1 0 0 0]
[ 2.  1.  7.  6.  2.  2.  9.  1.  9.  3.] [1 0 1 0]

当前模型采用sigmoid激活函数、binary cross-entropy损失函数与adam优化器,能识别出多标签中的至少一个标签,但无法同时识别多个标签。请问如何基于CNN或RNN实现多标签的准确识别?

现有实现代码

X_train, X_test, y_train, y_test = train_test_split(values, labels, test_size=0.2, random_state=42)
X_train = X_train.reshape((X_train.shape[0], X_train.shape[1], 1))
X_test = X_test.reshape((X_test.shape[0], X_test.shape[1], 1))

num_classes = len(np.unique(np.concatenate((y_train, y_test), axis=0)))
input_shape = X_train.shape[1:]

enc = sklearn.preprocessing.OneHotEncoder(categories='auto')
enc.fit(np.concatenate((y_train, y_test), axis=0).reshape(-1, 1))
y_train = enc.transform(y_train.reshape(-1, 1)).toarray()
y_test = enc.transform(y_test.reshape(-1, 1)).toarray()

#some model...

output_layer = Dense(units=num_classes, activation='sigmoid')(dense_layer)
model = Model(inputs=input_layer, outputs=output_layer)

model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])

batch_size = 32
nb_epochs = 500

if earlystopping:
 
    early_stopping = EarlyStopping(monitor='val_loss', patience=50, restore_best_weights=True)
 
    history = model.fit(X_train, y_train, batch_size=batch_size, epochs=nb_epochs,
                    validation_data=(X_test, y_test), callbacks=[early_stopping])

predictions = model.predict(test_values) #test_values are different than X_test.
解决方案

1. 修正数据预处理逻辑

你当前用OneHotEncoder处理标签的逻辑错误——多标签任务的标签本身就是4维二进制向量(单标签仅1位为1,多标签多位为1),不需要再做独热编码。现有编码方式会把每个标签位当成独立类别,完全破坏多标签结构。

修正步骤:

  • 确认训练集标签是4维二进制向量格式
  • 移除OneHotEncoder相关代码,直接使用原始标签矩阵

2. 适配多标签的CNN/RNN模型实现

基于CNN的模型(捕捉局部时域特征)

from tensorflow.keras.layers import Input, Conv1D, MaxPooling1D, Flatten, Dense, Dropout
from tensorflow.keras.models import Model

input_layer = Input(shape=input_shape)
# 第一层卷积:提取局部特征
x = Conv1D(filters=32, kernel_size=3, activation='relu')(input_layer)
x = MaxPooling1D(pool_size=2)(x)
# 第二层卷积:提取高阶抽象特征
x = Conv1D(filters=64, kernel_size=3, activation='relu')(x)
x = MaxPooling1D(pool_size=2)(x)
# 扁平化后接入全连接层
x = Flatten()(x)
x = Dense(64, activation='relu')(x)
x = Dropout(0.5)(x)  # 抑制过拟合
# 输出层:4个神经元对应4类,sigmoid激活实现多标签独立判定
output_layer = Dense(4, activation='sigmoid')(x)

model = Model(inputs=input_layer, outputs=output_layer)

基于RNN的模型(捕捉时序依赖特征)

from tensorflow.keras.layers import Input, LSTM, Dense, Dropout
from tensorflow.keras.models import Model

input_layer = Input(shape=input_shape)
# LSTM层:捕捉序列长期依赖关系
x = LSTM(64, return_sequences=True)(input_layer)
x = LSTM(32)(x)
x = Dense(32, activation='relu')(x)
x = Dropout(0.5)(x)
# 输出层保持sigmoid激活
output_layer = Dense(4, activation='sigmoid')(x)

model = Model(inputs=input_layer, outputs=output_layer)

3. 训练与推理优化

  • 调整评估指标:默认accuracy不适合多标签任务,替换为binary_accuracy或F1Score:
    from tensorflow.keras.metrics import BinaryAccuracy, F1Score
    model.compile(
        loss='binary_crossentropy',
        optimizer='adam',
        metrics=[BinaryAccuracy(name='binary_acc'), F1Score(average='macro', name='macro_f1')]
    )
    
  • 时间序列数据增强:对训练集做随机裁剪、加高斯噪声、时间翻转(任务允许时),提升模型泛化能力,帮助模型学习多标签特征组合。
  • 自定义推理阈值:默认0.5阈值可能不适合所有类别,可根据验证集F1分数调整,比如针对难识别类别降低阈值:
    # 针对4类分别设置阈值
    thresholds = np.array([0.4, 0.5, 0.45, 0.5])
    predictions = model.predict(test_values)
    predicted_labels = (predictions >= thresholds).astype(int)
    
  • 补充多标签训练样本:若条件允许,人工合成少量多标签样本(比如混合两类单标签样本特征),让模型学习“同时属于多类”的特征模式。

内容的提问来源于stack exchange,提问作者UtkE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 12:44:55