Keras timeseries_dataset_from_array是否支持多数据类型?处理方案咨询
使用
timeseries_dataset_from_array()处理多类型时序数据的问题与解决方案 问题核心:timeseries_dataset_from_array()是否仅支持单一数据类型?
是的,该函数要求输入的特征数据必须是单一数据类型的张量或数组。当你传入混合类型的Pandas DataFrame时,NumPy会自动将其转换为object类型的异构数组,而TensorFlow无法将这种数组转换为统一类型的张量,这就是触发ValueError的原因。官方文档虽未明确标注此限制,但函数底层实现依赖于输入数据的类型一致性。
解决方案:创建多类型时序训练序列的方法
方法1:拆分特征后分别生成时序数据集,再合并
将不同类型的特征拆分,各自生成时序数据集,最后用tf.data.Dataset.zip()合并。这种方法简单直接,且能保留各特征的原始类型,方便后续接入Keras预处理层。
import pandas as pd import numpy as np import tensorflow as tf from tensorflow.keras.utils import timeseries_dataset_from_array # 示例数据 X = pd.DataFrame({ "categorical": ["a", "b", "c", "a", "b"], "numerical": [1, 2, 3, 4, 5] }) y = np.array([1,2,3,4,5]) n_timesteps = 2 batch_size = 1 # 拆分分类与数值特征 cat_features = X["categorical"].values.reshape(-1, 1) num_features = X["numerical"].values.reshape(-1, 1) # 分别生成时序数据集(标签需对应时序窗口的目标值) cat_ds = timeseries_dataset_from_array( cat_features, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size ) num_ds = timeseries_dataset_from_array( num_features, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size ) y_ds = timeseries_dataset_from_array( y, None, sequence_length=n_timesteps, sequence_stride=1, batch_size=batch_size ) # 合并特征与标签 combined_ds = tf.data.Dataset.zip(((cat_ds, num_ds), y_ds)) # 验证输出 for (cat_seq, num_seq), y_seq in combined_ds: print("分类特征序列:\n", cat_seq) print("数值特征序列:\n", num_seq) print("标签序列:\n", y_seq)
方法2:用tf.data.Dataset手动构建时序窗口
直接使用TensorFlow的Dataset API构建时序窗口,天然支持结构化多类型特征,灵活性更强,适合复杂的时序需求。
import pandas as pd import numpy as np import tensorflow as tf # 示例数据 X = pd.DataFrame({ "categorical": ["a", "b", "c", "a", "b"], "numerical": [1, 2, 3, 4, 5] }) y = np.array([1,2,3,4,5]) n_timesteps = 2 batch_size = 1 # 将特征转换为TensorFlow结构化数据 features = { "categorical": tf.convert_to_tensor(X["categorical"], dtype=tf.string), "numerical": tf.convert_to_tensor(X["numerical"], dtype=tf.float32) } # 创建基础数据集 base_ds = tf.data.Dataset.from_tensor_slices((features, y)) # 定义生成时序窗口的函数 def create_windowed_dataset(ds, window_size): # 生成滑动窗口 window_ds = ds.window(window_size, shift=1, drop_remainder=True) # 展平窗口数据,提取特征序列与对应标签(取窗口最后一个标签作为目标) def flatten_window(window): feat_window, label_window = tf.data.Dataset.zip(window) return ( tf.data.Dataset.zip(feat_window).batch(window_size), label_window.batch(window_size).map(lambda x: x[-1]) ) return window_ds.flat_map(flatten_window) # 生成最终时序数据集 windowed_ds = create_windowed_dataset(base_ds, n_timesteps).batch(batch_size) # 验证输出 for feat_seq, y_seq in windowed_ds: print("分类特征序列:\n", feat_seq["categorical"]) print("数值特征序列:\n", feat_seq["numerical"]) print("目标标签:\n", y_seq)
方法3:提前预处理统一数据类型(备选)
如果允许提前对分类特征编码(如标签编码、独热编码),将所有特征转为数值类型后,即可直接使用timeseries_dataset_from_array()。但此方法不符合“在神经网络内完成预处理”的需求,仅作为场景受限的备选方案。
后续接入Keras预处理层的示例
以方法2的数据集为例,可直接在模型中对不同特征分支做预处理:
from tensorflow.keras import layers, Model # 分类特征分支 cat_input = layers.Input(shape=(n_timesteps,), dtype=tf.string, name="categorical") cat_processed = layers.StringLookup(vocabulary=["a", "b", "c"])(cat_input) cat_processed = layers.CategoryEncoding(num_tokens=3, output_mode="one_hot")(cat_processed) # 数值特征分支 num_input = layers.Input(shape=(n_timesteps,), dtype=tf.float32, name="numerical") num_processed = layers.Normalization()(num_input) # 合并特征并接入LSTM combined = layers.concatenate([cat_processed, num_processed]) lstm_out = layers.LSTM(32)(combined) output = layers.Dense(1)(lstm_out) # 构建模型 model = Model(inputs=[cat_input, num_input], outputs=output) model.compile(optimizer="adam", loss="mse") # 训练模型 model.fit(windowed_ds, epochs=5)
内容的提问来源于stack exchange,提问作者Requin
相关产品推荐
相关产品推荐

