You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas DataFrame的TF/Keras变长字符串序列特征处理问询

处理变长字符串列表序列特征的TF/Keras数据预处理方案

你的问题核心是变长序列特征的张量转换——Pandas里的变长字符串列表无法直接转为标准张量,必须先转为RaggedTensor(TensorFlow专门处理变长数据的结构),再用StringLookup做编码,之后才能对接LSTM层。以下是完整的处理流程和代码示例:

1. 构造样例数据

先模拟你的数据结构(序列特征为变长字符串列表,同时包含数值/分类标量特征):

import pandas as pd
import tensorflow as tf
from tensorflow.keras import layers

df = pd.DataFrame({
    "seq_feature": [["apple", "banana"], ["orange"], ["banana", "apple", "orange"], ["orange", "banana"]],
    "scalar_num": [1.2, 3.4, 5.6, 7.8],
    "scalar_cat": ["red", "blue", "red", "green"]
})

2. 处理序列特征

2.1 转为RaggedTensor

Pandas的列表列本质是object类型,无法直接转为标准张量,必须先转为RaggedTensor适配变长结构:

# 提取列表数据,生成RaggedTensor
seq_ragged = tf.ragged.constant(df["seq_feature"].tolist())

2.2 用StringLookup编码分类字符串

StringLookup可以直接处理RaggedTensor,无需先转成普通字符串张量:

# 初始化编码层,自动从数据中学习词汇表
string_lookup = layers.StringLookup(output_mode="int")
string_lookup.adapt(seq_ragged)

# 将字符串序列转为整数编码的RaggedTensor
seq_encoded = string_lookup(seq_ragged)

2.3 可选:序列填充(转为固定长度张量)

如果你的LSTM需要固定长度输入,用Padding层做后填充:

# 自动获取最大序列长度,或手动指定
max_seq_len = seq_encoded.bounding_shape()[1]
seq_padded = layers.Padding(padding="post", input_shape=(None,))(seq_encoded)

若不需要固定长度,TensorFlow的LSTM层可直接接收RaggedTensor,跳过此步骤即可。

3. 处理标量特征

3.1 数值标量

直接转为浮点型张量:

scalar_num_tensor = tf.convert_to_tensor(df["scalar_num"], dtype=tf.float32)

3.2 分类标量

同样用StringLookup编码,可选转为one-hot格式:

scalar_cat_lookup = layers.StringLookup(output_mode="one_hot")
scalar_cat_lookup.adapt(df["scalar_cat"])
scalar_cat_encoded = scalar_cat_lookup(df["scalar_cat"])

4. 构建融合模型

分别将序列特征输入LSTM、标量特征输入MLP,最后用全连接层融合输出:

# 序列分支:支持RaggedTensor输入
seq_input = layers.Input(shape=(None,), dtype=tf.int32, ragged=True)
x_seq = layers.Embedding(input_dim=string_lookup.vocabulary_size(), output_dim=32)(seq_input)
x_seq = layers.LSTM(64)(x_seq)

# 标量分支:合并数值与分类特征
scalar_num_input = layers.Input(shape=(1,))
scalar_cat_input = layers.Input(shape=(scalar_cat_lookup.vocabulary_size(),))
x_scalar = layers.concatenate([scalar_num_input, scalar_cat_input])
x_scalar = layers.Dense(32, activation="relu")(x_scalar)
x_scalar = layers.Dense(16, activation="relu")(x_scalar)

# 融合输出
merged = layers.concatenate([x_seq, x_scalar])
output = layers.Dense(1, activation="sigmoid")(merged)

model = tf.keras.Model(
    inputs=[seq_input, scalar_num_input, scalar_cat_input],
    outputs=output
)

model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])

5. 用Dataset管道喂入数据

为提升训练效率,建议用tf.data.Dataset包装处理好的数据:

# 构造数据集(需替换为你的实际标签列)
dataset = tf.data.Dataset.from_tensor_slices({
    "seq_input": seq_encoded,
    "scalar_num_input": scalar_num_tensor,
    "scalar_cat_input": scalar_cat_encoded
})
# 假设标签列为"label",添加标签映射
# dataset = dataset.map(lambda x: (x, x["label"]))
dataset = dataset.batch(2)

# 启动训练
# model.fit(dataset, epochs=10)

你之前报错的原因

直接用StringLookup处理Pandas的列表列时,喂入的是Python列表数组而非TensorFlow的张量结构,变长列表无法转为标准张量导致报错。正确顺序是先转RaggedTensor,再做编码,才能生成LSTM可识别的张量格式。

内容的提问来源于stack exchange,提问作者Douglas Chien

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:45:27