如何用tf.data构建多变量时间序列数据集,解决LSTM多特征输入报错问题
问题原因
当Dataset的输出是多元素元组时,window()方法会对每个位置的元素分别生成独立的窗口数据集,因此flat_map调用映射函数时,会传入和元组元素数量相等的参数,你仅定义了1个入参的sub_to_batch,因此触发参数数量不匹配的报错。
解决代码
方案1:调整窗口处理函数适配多特征
import tensorflow as tf class generator: def __init__(self, n=5): self.n = n def __call__(self): for i in range(self.n): yield (i, 10*i) dataset = tf.data.Dataset.from_generator( generator(), output_signature=( tf.TensorSpec(shape=(), dtype=tf.uint16), tf.TensorSpec(shape=(), dtype=tf.int32) ) ) window_size = 3 windows = dataset.window(window_size, shift=1) # 接收两个参数对应两个特征的窗口 def sub_to_batch(sub1, sub2): # 分别对两个特征做batch b1 = sub1.batch(window_size, drop_remainder=True) b2 = sub2.batch(window_size, drop_remainder=True) # 沿最后一维拼接两个特征,得到(窗口大小, 2)的结构 return tf.data.Dataset.zip((b1, b2)).map(lambda x, y: tf.stack([x, y], axis=-1)) final_dset = windows.flat_map(sub_to_batch) print(list(final_dset.as_numpy_iterator()))
方案2:提前拼接特征简化后续处理
你也可以在创建窗口前先把多特征元组拼接为单个张量,后续处理逻辑和单特征场景完全一致:
import tensorflow as tf class generator: def __init__(self, n=5): self.n = n def __call__(self): for i in range(self.n): yield (i, 10*i) dataset = tf.data.Dataset.from_generator( generator(), output_signature=( tf.TensorSpec(shape=(), dtype=tf.uint16), tf.TensorSpec(shape=(), dtype=tf.int32) ) ) # 提前把两个特征拼接为单个张量,shape为(2,) dataset = dataset.map(lambda x, y: tf.stack([x, y], axis=-1)) window_size = 3 windows = dataset.window(window_size, shift=1) def sub_to_batch(sub): return sub.batch(window_size, drop_remainder=True) final_dset = windows.flat_map(sub_to_batch) print(list(final_dset.as_numpy_iterator()))
输出结果
两种方案运行后都能得到你预期的输出:
[ array([[ 0, 0], [ 1, 10], [ 2, 20]], dtype=int32), array([[ 1, 10], [ 2, 20], [ 3, 30]], dtype=int32), array([[ 2, 20], [ 3, 30], [ 4, 40]], dtype=int32) ]
整体shape为(3, 3, 2),可直接作为LSTM模型的输入。
内容的提问来源于stack exchange,提问作者DocDriven
相关产品推荐
相关产品推荐

