You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras timeseries_dataset_from_array是否支持多数据类型?处理方案咨询

使用timeseries_dataset_from_array()处理多类型时序数据的问题与解决方案

问题核心:timeseries_dataset_from_array()是否仅支持单一数据类型?

是的,该函数要求输入的特征数据必须是单一数据类型的张量或数组。当你传入混合类型的Pandas DataFrame时,NumPy会自动将其转换为object类型的异构数组,而TensorFlow无法将这种数组转换为统一类型的张量,这就是触发ValueError的原因。官方文档虽未明确标注此限制,但函数底层实现依赖于输入数据的类型一致性。

解决方案:创建多类型时序训练序列的方法

方法1:拆分特征后分别生成时序数据集,再合并

将不同类型的特征拆分,各自生成时序数据集,最后用tf.data.Dataset.zip()合并。这种方法简单直接,且能保留各特征的原始类型,方便后续接入Keras预处理层。

import pandas as pd
import numpy as np
import tensorflow as tf
from tensorflow.keras.utils import timeseries_dataset_from_array

# 示例数据
X = pd.DataFrame({
    "categorical": ["a", "b", "c", "a", "b"],
    "numerical": [1, 2, 3, 4, 5]
})
y = np.array([1,2,3,4,5])
n_timesteps = 2
batch_size = 1

# 拆分分类与数值特征
cat_features = X["categorical"].values.reshape(-1, 1)
num_features = X["numerical"].values.reshape(-1, 1)

# 分别生成时序数据集(标签需对应时序窗口的目标值)
cat_ds = timeseries_dataset_from_array(
    cat_features, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size
)
num_ds = timeseries_dataset_from_array(
    num_features, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size
)
y_ds = timeseries_dataset_from_array(
    y, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size
)

# 合并特征与标签
combined_ds = tf.data.Dataset.zip(((cat_ds, num_ds), y_ds))

# 验证输出
for (cat_seq, num_seq), y_seq in combined_ds:
    print("分类特征序列:\n", cat_seq)
    print("数值特征序列:\n", num_seq)
    print("标签序列:\n", y_seq)

方法2:用tf.data.Dataset手动构建时序窗口

直接使用TensorFlow的Dataset API构建时序窗口,天然支持结构化多类型特征,灵活性更强,适合复杂的时序需求。

import pandas as pd
import numpy as np
import tensorflow as tf

# 示例数据
X = pd.DataFrame({
    "categorical": ["a", "b", "c", "a", "b"],
    "numerical": [1, 2, 3, 4, 5]
})
y = np.array([1,2,3,4,5])
n_timesteps = 2
batch_size = 1

# 将特征转换为TensorFlow结构化数据
features = {
    "categorical": tf.convert_to_tensor(X["categorical"], dtype=tf.string),
    "numerical": tf.convert_to_tensor(X["numerical"], dtype=tf.float32)
}

# 创建基础数据集
base_ds = tf.data.Dataset.from_tensor_slices((features, y))

# 定义生成时序窗口的函数
def create_windowed_dataset(ds, window_size):
    # 生成滑动窗口
    window_ds = ds.window(window_size, shift=1, drop_remainder=True)
    # 展平窗口数据,提取特征序列与对应标签(取窗口最后一个标签作为目标)
    def flatten_window(window):
        feat_window, label_window = tf.data.Dataset.zip(window)
        return (
            tf.data.Dataset.zip(feat_window).batch(window_size),
            label_window.batch(window_size).map(lambda x: x[-1])
        )
    return window_ds.flat_map(flatten_window)

# 生成最终时序数据集
windowed_ds = create_windowed_dataset(base_ds, n_timesteps).batch(batch_size)

# 验证输出
for feat_seq, y_seq in windowed_ds:
    print("分类特征序列:\n", feat_seq["categorical"])
    print("数值特征序列:\n", feat_seq["numerical"])
    print("目标标签:\n", y_seq)

方法3:提前预处理统一数据类型(备选)

如果允许提前对分类特征编码(如标签编码、独热编码),将所有特征转为数值类型后,即可直接使用timeseries_dataset_from_array()。但此方法不符合“在神经网络内完成预处理”的需求,仅作为场景受限的备选方案。

后续接入Keras预处理层的示例

以方法2的数据集为例,可直接在模型中对不同特征分支做预处理:

from tensorflow.keras import layers, Model

# 分类特征分支
cat_input = layers.Input(shape=(n_timesteps,), dtype=tf.string, name="categorical")
cat_processed = layers.StringLookup(vocabulary=["a", "b", "c"])(cat_input)
cat_processed = layers.CategoryEncoding(num_tokens=3, output_mode="one_hot")(cat_processed)

# 数值特征分支
num_input = layers.Input(shape=(n_timesteps,), dtype=tf.float32, name="numerical")
num_processed = layers.Normalization()(num_input)

# 合并特征并接入LSTM
combined = layers.concatenate([cat_processed, num_processed])
lstm_out = layers.LSTM(32)(combined)
output = layers.Dense(1)(lstm_out)

# 构建模型
model = Model(inputs=[cat_input, num_input], outputs=output)
model.compile(optimizer="adam", loss="mse")

# 训练模型
model.fit(windowed_ds, epochs=5)

内容的提问来源于stack exchange,提问作者Requin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 20:50:29