You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Keras的Conv1D正确重塑一维输入数据?

问题描述

我的数据集包含68个样本、10000个特征,第一列为年龄标签(训练时需忽略)。当前用Conv1D模型训练的性能很差,我猜测问题出在数据编码/重塑方式上——因为数据是一维结构,很难区分各特征的维度对应关系,现在想知道Conv1D的正确数据重塑方法。

数据示例

30, 0.5, 0.2, 0.004, 0.001, 0.1, 0.003, 0.0005, 0.003
20, 0.1, 0.003, 0.0005, 0.003, 0.003, 0.1, 0.4, 0.33
25, 0.9, 0.63, 0.0005, 0.003, 0.0005, 0.003, 0.1, 0.003
26, 0.08, 0.83, 0.0005, 0.003, 0.1, 0.003, 0.0005, 0.003
39, 0.003, 0.1, 0.4, 0.33, 0.9, 0.63, 0.0005, 0.003

(第一列为年龄,其余为0-1之间的数值特征)

数据读取与划分代码

import pandas as pd
from sklearn.model_selection import train_test_split

data1_df = pd.read_csv("GSE106648_data1.csv")
data2_df = pd.read_csv("GSE106648_data2.csv")

# Split the data
X1, y1 = data1_df.values[:,1:], data1_df.values[:,0]
X2, y2 = data2_df.values[:,1:], data2_df.values[:,0]

X1_train, X1_valid, y1_train, y1_valid = train_test_split(X1, y1, test_size=0.2, shuffle= True)
X2_train, X2_valid, y2_train, y2_valid = train_test_split(X2, y2, test_size=0.2, shuffle= True)

当前数据重塑方式

sample_size = X1_train.shape[0] # number of samples in train set
time_steps  = X1_train.shape[1] # number of features in train set
input_dimension = 1               # each feature is represented by 1 number

# We need to reshape the Test and validation data as well:
X1_train_reshaped = X1_train.reshape(X1_train.shape[0],X1_train.shape[1],1)
X1_valid_reshaped = X1_valid.reshape(X1_valid.shape[0],X1_valid.shape[1],1)
X2_train_reshaped = X2_train.reshape(X2_train.shape[0],X2_train.shape[1],1)
X2_valid_reshaped = X2_valid.reshape(X2_valid.shape[0],X2_valid.shape[1],1)
X1_reshaped = X1.reshape(X1.shape[0],X1.shape[1],1)
X2_reshaped = X2.reshape(X2.shape[0],X2.shape[1],1)

当前Conv1D模型代码

from tensorflow.keras import Sequential, layers, optimizers

def conv1D_model():
    n_timesteps = X1_train_reshaped.shape[1] 
    n_features  = X1_train_reshaped.shape[2] #1 
    
    model = Sequential(name="model_conv1D")
    model.add(layers.Input(shape=(n_timesteps,n_features)))    
    model.add(layers.Conv1D(filters=8, kernel_size=4, activation='LeakyReLU', name="Conv1D_1"))
    model.add(layers.MaxPooling1D(pool_size=4, name="MaxPooling1D_1"))
    model.add(layers.Flatten())
    model.add(layers.Dense(50, activation='LeakyReLU', name="Dense_1"))
    model.add(layers.Dropout(0.5))
    model.add(layers.Dense(n_features, activation='LeakyReLU', name="output"))
    
    optimizer = optimizers.Adam(learning_rate=1e-4)

    model.compile(loss='mse',optimizer=optimizer,metrics=['mae'])
    return model

my_model = conv1D_model()
history = my_model.fit(X1_train_reshaped,y1_train,batch_size=50,epochs=100,validation_data=(X1_valid_reshaped,y1_valid), shuffle=True)

问题分析与解决方案

1. 数据重塑的核心判断

你当前的(样本数, 特征数, 1)重塑格式本身符合Conv1D的输入要求(Conv1D期望输入形状为(samples, timesteps, features)),但性能差的根源不在重塑本身,而是两个关键问题:特征是否具备Conv1D能利用的序列关联性,以及模型结构的明显错误。

关键前提:特征是否有序?

Conv1D的设计目标是捕捉有序序列中的局部模式(比如时间序列、基因位点序列、文本序列)。如果你的10000个特征是无序的(比如随机排列的基因表达量、无关联的特征集合),Conv1D的滑动窗口无法学到有效模式,此时用全连接网络反而更合适。

如果特征是有序的(比如按基因组位置排列的位点、时序采集的信号),当前重塑方式是正确的,只需优化模型结构。

2. 模型结构的致命错误

你的输出层设置完全不符合回归任务需求:

model.add(layers.Dense(n_features, activation='LeakyReLU', name="output"))

你要预测的是单值年龄(回归任务),输出层应设为Dense(1),且不需要激活函数(年龄是连续值,LeakyReLU会限制输出为正数,约束模型拟合能力)。修正后的输出层:

model.add(layers.Dense(1, name="output"))  # 回归任务无需激活函数

3. 高维特征的Conv1D优化建议

针对10000个高维特征,当前模型参数设置过于保守,容易导致欠拟合或过拟合:

  • 增加卷积核数量:将filters=8改为32或64,让模型捕捉更多局部模式
  • 优化池化策略:用GlobalAveragePooling1D替代MaxPooling1D + Flatten,大幅减少参数数量,避免过拟合:
    model.add(layers.Conv1D(filters=32, kernel_size=4, activation='LeakyReLU', name="Conv1D_1"))
    model.add(layers.GlobalAveragePooling1D(name="GlobalAvgPool_1"))  # 直接输出每个过滤器的均值
    

4. 特征分组重塑(可选,需基于特征含义)

如果你的10000个特征是由明确子维度组成的(比如100个维度×100个特征),可将数据重塑为(样本数, 子维度长度, 子维度数量),让Conv1D在子维度内部捕捉模式。例如:

# 假设特征是100个时序点×100个传感器的组合
X1_train_reshaped = X1_train.reshape(X1_train.shape[0], 100, 100)

注意:此操作必须基于对特征含义的明确认知,不能随意分组。

5. 数据预处理优化

回归任务中对标签做标准化处理,能大幅提升模型收敛速度:

from sklearn.preprocessing import StandardScaler

scaler_y = StandardScaler()
y1_train_scaled = scaler_y.fit_transform(y1_train.reshape(-1,1))
y1_valid_scaled = scaler_y.transform(y1_valid.reshape(-1,1))

# 训练时用缩放后的标签,预测后再逆变换得到真实年龄
history = my_model.fit(X1_train_reshaped, y1_train_scaled, batch_size=50, epochs=100, validation_data=(X1_valid_reshaped, y1_valid_scaled), shuffle=True)

内容的提问来源于stack exchange,提问作者Caterina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 14:30:16