You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Keras实现文本(Embedding+LSTM)与数值特征融合的分类模型?

嘿,这个需求在Keras里用Functional API就能完美解决——这是处理多输入特征模型的标准方案,我给你捋清楚每一步,附上可直接运行的代码示例:

1. 先搞定数据预处理

因为文本特征和数值特征的处理逻辑完全不同,得分开预处理:

文本特征处理

首先要把文本转成模型能识别的序列格式,用Keras自带的工具就能搞定:

from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences

# 假设texts是你的文本特征列表(比如每条评论内容)
tokenizer = Tokenizer(num_words=10000)  # 只保留前10000个高频词,可按需调整
tokenizer.fit_on_texts(texts)
text_sequences = tokenizer.texts_to_sequences(texts)
# 统一序列长度,太长截断、太短补0,maxlen根据你的文本平均长度调整
text_padded = pad_sequences(text_sequences, maxlen=100)

数值特征处理

数值特征(年龄、类别编码等)需要做标准化/归一化,避免不同尺度的特征影响模型训练:

from sklearn.preprocessing import StandardScaler
import numpy as np

# 假设numeric_features是形状为(样本数, 数值特征数量)的二维数组
scaler = StandardScaler()
numeric_scaled = scaler.fit_transform(numeric_features)
2. 构建双分支模型

用Functional API分别搭建文本分支和数值分支,最后把两个分支的输出拼接起来:

from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Embedding, LSTM, Dense, concatenate, Dropout

# ---------------------- 文本分支 ----------------------
# 定义文本输入层,shape对应预处理后的序列长度
text_input = Input(shape=(100,), name="text_input")
# Embedding层:把词序列转成词向量,input_dim要和Tokenizer的num_words一致
embedding = Embedding(input_dim=10000, output_dim=128, input_length=100)(text_input)
# LSTM层:提取文本的时序特征,return_sequences=False表示只取最后一个时间步的输出
lstm_out = LSTM(64, return_sequences=False)(embedding)
# 加Dropout防止过拟合
text_branch = Dropout(0.3)(lstm_out)

# ---------------------- 数值分支 ----------------------
# 定义数值输入层,shape对应数值特征的数量
numeric_input = Input(shape=(numeric_scaled.shape[1],), name="numeric_input")
# 用Dense层提取数值特征的非线性关系
dense1 = Dense(32, activation="relu")(numeric_input)
dense2 = Dense(16, activation="relu")(dense1)
numeric_branch = Dropout(0.3)(dense2)

# ---------------------- 拼接分支+输出 ----------------------
# 把两个分支的输出拼接在一起
concat_layer = concatenate([text_branch, numeric_branch])
# 再加一层Dense做特征融合
final_dense = Dense(8, activation="relu")(concat_layer)
# 最终输出层:二分类用sigmoid,多分类用softmax
output = Dense(1, activation="sigmoid")(final_dense)

# 定义完整模型,指定输入和输出
model = Model(inputs=[text_input, numeric_input], outputs=output)
3. 编译与训练模型

最后就是编译模型、喂数据训练了:

# 编译模型:二分类用binary_crossentropy,多分类换categorical_crossentropy
model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])

# 训练时要同时传入两个预处理后的输入特征
history = model.fit(
    [text_padded, numeric_scaled],
    labels,  # 你的标签数组(二分类是0/1,多分类是整数或one-hot编码)
    epochs=10,
    batch_size=32,
    validation_split=0.2  # 用20%数据做验证
)
几个关键注意点
  • 文本序列的maxlen:要根据你的文本长度分布调整,比如先统计所有文本的长度,取95%分位数作为maxlen,避免浪费计算资源或丢失关键信息;
  • 多分类适配:如果是多分类任务,把输出层的激活函数改成softmax,损失函数用sparse_categorical_crossentropy(标签是整数)或categorical_crossentropy(标签是one-hot编码);
  • 调参空间:可以根据样本量调整LSTM、Dense层的神经元数量,以及Dropout的比例,避免过拟合;
  • 多文本特征:如果有多个文本特征(比如标题+正文),可以再加一个文本分支,最后一起拼接进concat层。

内容的提问来源于stack exchange,提问作者Sreeram TP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:11:33