You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Keras模型输入多特征数据集并适配简单CNN模型?

嘿,作为Keras新手碰到这种混合特征的问题太正常了,别担心!我一步步给你拆解怎么把这个数据集接到CNN模型里~

首先搞定数据预处理(这是让模型“看懂”数据的核心!)

你的数据有文本、数值、多分类标签三种类型,得分别处理:

1. 处理文本特征(Subject + message)

CNN擅长处理序列数据,但文字得先转成模型能理解的数值序列:

  • 先把Subject和message合并成一个字段(比如combined_text = df['Subject'] + ' ' + df['message']),这样模型能同时学习标题和内容的信息;
  • 用Keras的Tokenizer来分词、转序列:
    from tensorflow.keras.preprocessing.text import Tokenizer
    from tensorflow.keras.preprocessing.sequence import pad_sequences
    
    # 初始化Tokenizer,设置最大词汇量(可以根据你的数据调整)
    tokenizer = Tokenizer(num_words=10000, oov_token="<OOV>")
    tokenizer.fit_on_texts(df['combined_text'])
    
    # 把文本转成数字序列
    text_sequences = tokenizer.texts_to_sequences(df['combined_text'])
    
    # 统一序列长度(CNN需要固定输入尺寸,太短补0,太长截断)
    max_seq_len = 200  # 可以根据你的文本平均长度调整
    padded_sequences = pad_sequences(text_sequences, maxlen=max_seq_len, padding='post', truncating='post')
    
  • 要是你数据量不大,用预训练的词嵌入(比如GloVe)能让模型效果更好,后面模型部分我会提怎么加。

2. 处理数值特征(origin_date + is_logged_in + use_cate)

这些数值特征得先做标准化,避免尺度差异影响模型:

from sklearn.preprocessing import StandardScaler

# 提取数值特征列
num_features = df[['origin_date', 'is_logged_in', 'use_cate']]
# 注意:如果origin_date是日期格式,得先转成数值!比如转成时间戳:
# df['origin_date'] = pd.to_datetime(df['origin_date']).astype('int64') // 10**9

scaler = StandardScaler()
scaled_num_features = scaler.fit_transform(num_features)

3. 处理41类的目标变量

字符串标签得转成独热编码,适合多分类任务:

from tensorflow.keras.utils import to_categorical
from sklearn.preprocessing import LabelEncoder

# 先把字符串标签转成整数ID
label_encoder = LabelEncoder()
integer_labels = label_encoder.fit_transform(df['target_variable'])

# 转成独热编码(41类就会变成41维的向量)
one_hot_labels = to_categorical(integer_labels, num_classes=41)
然后构建多输入的CNN模型

因为咱们有文本和数值两种输入,得用Keras的函数式API来把两个分支的输出融合起来:

from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Embedding, Conv1D, GlobalMaxPooling1D, Dense, concatenate, Dropout

# 第一个分支:处理文本序列
text_input = Input(shape=(max_seq_len,), name='text_input')
# 嵌入层:把数字序列转成稠密向量
embedding_layer = Embedding(input_dim=10000, output_dim=128, input_length=max_seq_len)(text_input)
# CNN层提取文本特征
conv_layer = Conv1D(filters=64, kernel_size=3, activation='relu')(embedding_layer)
# 全局最大池化,把序列转成固定长度的特征向量
pool_layer = GlobalMaxPooling1D()(conv_layer)
text_output = Dense(64, activation='relu')(pool_layer)

# 第二个分支:处理数值特征
num_input = Input(shape=(3,), name='num_input')  # 3个数值特征,所以输入维度是3
num_output = Dense(32, activation='relu')(num_input)

# 把两个分支的特征融合在一起
combined_features = concatenate([text_output, num_output])
# 加个Dropout防止过拟合
combined_dropout = Dropout(0.5)(combined_features)
# 最后输出41类的概率,用softmax激活
final_output = Dense(41, activation='softmax')(combined_dropout)

# 定义完整模型
model = Model(inputs=[text_input, num_input], outputs=final_output)
# 编译模型:多分类用categorical_crossentropy损失
model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
最后训练模型

把预处理好的数据拆分训练集和测试集,然后喂给模型就行:

from sklearn.model_selection import train_test_split

# 拆分数据集,测试集占20%
X_text_train, X_text_test, X_num_train, X_num_test, y_train, y_test = train_test_split(
    padded_sequences, scaled_num_features, one_hot_labels, test_size=0.2, random_state=42
)

# 开始训练
history = model.fit(
    [X_text_train, X_num_train], y_train,
    epochs=10,  # 可以根据验证集精度调整
    batch_size=32,
    validation_data=([X_text_test, X_num_test], y_test)
)
几个小Tips帮你优化
  • 要是文本数据少,替换嵌入层用预训练GloVe:先加载GloVe的词嵌入矩阵,然后把Embedding层改成Embedding(input_dim=10000, output_dim=100, weights=[embedding_matrix], input_length=max_seq_len, trainable=False)(output_dim要和GloVe的维度一致);
  • 超参数比如max_seq_len、CNN的filters、kernel_size都可以根据你的数据调整,比如文本长就把max_seq_len调大;
  • 训练时如果验证集精度下降,可能是过拟合,可以加更多Dropout或者减少训练轮数。

内容的提问来源于stack exchange,提问作者Aditya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:21:16