You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Keras的LSTM模型中融合文本与类别特征的实现方法

可行方法实现步骤

当然有可行的方法,这类融合文本特征与类别变量的分类任务在机器学习中十分常见,具体可以按以下流程实现:

1. 特征预处理

分别对Text(文本特征)和Topic(类别变量)做针对性处理,把它们转换成模型能识别的数值特征:

文本特征处理

  • 传统统计特征:用TF-IDF或词袋模型(CountVectorizer)将文本转换成稀疏数值矩阵,这类方法简单高效,适合小数据场景。例如用TfidfVectorizer可以提取文本中词语的重要性权重。
  • 预训练词嵌入/语义特征:如果追求更好的语义理解,可以用Word2Vec、BERT等预训练模型提取文本的语义向量,把每个文本转换成固定维度的稠密向量,适合大数据或复杂语义场景。

类别变量处理

Topic是离散类别值(如-1、1、2),不能直接当作数值输入模型,需要做编码转换:

  • 独热编码:用pandas.get_dummies()或sklearn.preprocessing.OneHotEncoder把每个类别值转换成独立的二进制特征列,避免模型误认为类别间有数值大小关系。
  • 若Topic本身存在顺序意义(比如表示优先级),也可以用标签编码,但从你的数据看,独热编码更合适。

2. 特征合并

把处理好的文本特征和编码后的Topic特征拼接成完整的特征矩阵,比如用numpy.hstack()拼接数组,或pandas.concat()拼接DataFrame列。

3. 模型训练

根据数据规模和需求选择合适的模型:

传统机器学习模型

适合中小规模数据,流程简单易解释:
可以用Pipeline把预处理和模型串起来,避免数据泄露,比如逻辑回归、随机森林、XGBoost都是不错的选择。示例代码:

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

# 假设你的数据集是df
X = df[['Text', 'Topic']]
y = df['Label']

# 拆分训练测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 定义特征预处理流程
preprocessor = ColumnTransformer(
    transformers=[
        # 处理文本特征
        ('text_tfidf', TfidfVectorizer(stop_words='english'), 'Text'),
        # 处理Topic类别特征
        ('topic_onehot', OneHotEncoder(sparse_output=False, handle_unknown='ignore'), ['Topic'])
    ])

# 构建完整训练管道
model_pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', LogisticRegression())
])

# 训练模型
model_pipeline.fit(X_train, y_train)

# 预测
y_pred = model_pipeline.predict(X_test)

深度学习多输入模型

适合大规模数据,能更好地捕捉复杂特征交互:
可以构建多输入模型,一个分支处理文本序列,另一个分支处理编码后的Topic特征,最后合并两个分支的输出做分类。示例代码(用TensorFlow/Keras):

import tensorflow as tf
from tensorflow.keras.layers import Input, Embedding, LSTM, Dense, concatenate
from tensorflow.keras.models import Model
from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
import pandas as pd

# 文本预处理:转成序列并填充
tokenizer = Tokenizer(num_words=10000)
tokenizer.fit_on_texts(df['Text'])
text_sequences = tokenizer.texts_to_sequences(df['Text'])
text_padded = pad_sequences(text_sequences, maxlen=50)  # 统一文本长度为50

# Topic预处理:独热编码
topic_onehot = pd.get_dummies(df['Topic'], prefix='Topic').values

# 定义多输入结构
text_input = Input(shape=(50,), name='text_input')
topic_input = Input(shape=(topic_onehot.shape[1],), name='topic_input')

# 文本分支:Embedding+LSTM捕捉语义
text_embedding = Embedding(input_dim=10000, output_dim=64)(text_input)
text_feature = LSTM(32)(text_embedding)

# Topic分支:全连接层处理类别特征
topic_feature = Dense(16, activation='relu')(topic_input)

# 合并两个分支的特征
merged_feature = concatenate([text_feature, topic_feature])
output = Dense(1, activation='sigmoid')(merged_feature)  # 二分类输出

# 构建并编译模型
model = Model(inputs=[text_input, topic_input], outputs=output)
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])

# 训练模型
model.fit([text_padded, topic_onehot], df['Label'], epochs=10, batch_size=32, validation_split=0.2)

额外注意点

  • 文本预处理时可以加入小写转换、标点去除、停用词过滤等步骤,提升特征质量。
  • 训练前要做数据拆分(训练/验证/测试集),避免过拟合。
  • 可以根据模型效果调整超参数,比如TF-IDF的ngram范围、LSTM的神经元数量等。

内容的提问来源于stack exchange,提问作者coelidonum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:50:22