合并异质数据TensorFlow数据集构建窗口数据集报错问题
多数据类型DataFrame构建TensorFlow窗口数据集修复方案
问题背景
需要对包含多种dtype的DataFrame构建窗口化时序数据集,此前同构数据的处理方案在传入字典结构的异构DataFrame时,直接使用flat_map会触发报错:AttributeError: 'dict' object has no attribute 'batch'。
不使用flat_map的初始实现会遇到两类核心问题:
- 类型错误:窗口生成后直接传入模型触发
TypeError: Inputs to a layer should be tensors. Got: <_VariantDataset element_spec=TensorSpec(shape=(), dtype=tf.string, name=None)>,本质是window方法返回的VariantDataset嵌套子数据集结构,没有被转换为模型可接收的张量格式 - 形状不匹配错误:尝试将嵌套结构转换为张量后,触发
StringLookup层报错:ValueError: Exception encountered when calling layer "string_lookup" (type StringLookup). When output_mode is not 'int', maximum supported output rank is 2. Received output_mode one_hot and input shape (None, None), which would result in output rank 3.,本质是窗口化后字符串特征增加了时间步维度,one_hot编码模式默认不支持3维输入
初始问题复现代码
测试数据构造
import pandas as pd import numpy as np import tensorflow as tf # 构造包含字符串、整数两种dtype的测试DataFrame x = pd.DataFrame({'col1': list('abcdefghij'), 'col2': np.arange(10), 'col3': np.arange(10)}) y = np.arange(10)
初始窗口数据集构建
window_size_x = 3 window_size_y = 2 shift_size = 1 x_slice = x[:-window_size_y] y_slice = y[window_size_x:] ds_x = tf.data.Dataset.from_tensor_slices(dict(x_slice)).window(window_size_x, shift=shift_size, drop_remainder=True) ds_y = tf.data.Dataset.from_tensor_slices(y_slice).window(window_size_y, shift=shift_size, drop_remainder=True) dataset = tf.data.Dataset.zip((ds_x, ds_y)) dataset = dataset.batch(1) # 打印单条样本,可见输出为嵌套VariantDataset结构而非张量 for i, j in dataset.take(1): print(i, j)
初始预处理器与模型代码
# 构造多输入预处理器 inputs = {'col1': tf.keras.Input(shape=(), name='col1', dtype=tf.string), 'col2': tf.keras.Input(shape=(), name='col2', dtype=tf.float32), 'col3': tf.keras.Input(shape=(), name='col3', dtype=tf.float32)} vocab = sorted(set(x_slice['col1'])) lookup = tf.keras.layers.StringLookup(vocabulary=vocab, output_mode='one_hot') lookup = lookup(inputs['col1'][:tf.newaxis]) numeric = tf.stack([tf.cast(inputs[i], dtype=tf.float32) for i in ['col2', 'col3']], axis=-1) preprocess_result = tf.concat([lookup, numeric], axis=-1) preprocessor = tf.keras.Model(inputs, preprocess_result) # 直接传入全量切片数据测试,预处理器可正常输出 preprocessor(dict(x_slice)) # 构造端到端模型并尝试训练 body = tf.keras.models.Sequential([tf.keras.layers.Dense(8), tf.keras.layers.Dense(window_size_y)]) x_processed = preprocessor(inputs) model_output = body(x_processed) model = tf.keras.Model(inputs, model_output) model.compile(loss='mae', optimizer='adam') # 执行训练触发前述报错 model.fit(dataset)
修复步骤
1. 修复VariantDataset无法转张量的问题
此前直接使用flat_map触发属性错误的原因,是flat_map要求返回的Dataset结构直接支持batch操作,传入字典时无法自动递归处理内部的VariantDataset,必须手动逐列对窗口子数据集做batch转换,将每个窗口聚合为固定形状的张量。
修复后的数据集构建代码:
window_size_x = 3 window_size_y = 2 shift_size = 1 x_slice = x[:-window_size_y] y_slice = y[window_size_x:] ds_x = tf.data.Dataset.from_tensor_slices(dict(x_slice)).window(window_size_x, shift=shift_size, drop_remainder=True) # 逐列处理字典内的特征,将每个窗口的子Dataset batch为形状(window_size_x,)的张量 ds_x = ds_x.map(lambda col_dict: {col_name: window_ds.batch(window_size_x) for col_name, window_ds in col_dict.items()}) ds_y = tf.data.Dataset.from_tensor_slices(y_slice).window(window_size_y, shift=shift_size, drop_remainder=True) # 标签窗口同步batch为张量 ds_y = ds_y.map(lambda window_ds: window_ds.batch(window_size_y)) dataset = tf.data.Dataset.zip((ds_x, ds_y)) # 验证单条样本输出,此时返回标准张量 for x_batch, y_batch in dataset.take(1): print({k: v.shape for k, v in x_batch.items()}, y_batch.shape)
运行验证输出:
{'col1': TensorShape([3]), 'col2': TensorShape([3]), 'col3': TensorShape([3])} TensorShape([2])
2. 修复StringLookup层3维输入报错
窗口化后单样本特征形状为(window_size_x,),比预处理器初始测试的单步输入多了时间步维度,需要调整输入层形状适配窗口结构,同时调整字符串编码逻辑:直接对形状为(batch_size, window_size_x)的字符串输入做one_hot编码,输出维度自动对齐为(batch_size, window_size_x, vocab_size+1),和扩维后的数值特征拼接即可。
修复后的预处理器与模型代码:
# 调整输入层形状,适配窗口化后(窗口长度,)的单样本特征维度 inputs = {'col1': tf.keras.Input(shape=(window_size_x,), name='col1', dtype=tf.string), 'col2': tf.keras.Input(shape=(window_size_x,), name='col2', dtype=tf.float32), 'col3': tf.keras.Input(shape=(window_size_x,), name='col3', dtype=tf.float32)} vocab = sorted(set(x_slice['col1'])) lookup_layer = tf.keras.layers.StringLookup(vocabulary=vocab, output_mode='one_hot') # 对字符串特征做one_hot编码,输出形状为(窗口长度, 词表大小+1) col1_encoded = lookup_layer(inputs['col1']) # 处理数值特征,对齐维度后和字符串编码结果拼接 col2_cast = tf.cast(tf.expand_dims(inputs['col2'], axis=-1), tf.float32) col3_cast = tf.cast(tf.expand_dims(inputs['col3'], axis=-1), tf.float32) preprocess_result = tf.concat([col1_encoded, col2_cast, col3_cast], axis=-1) preprocessor = tf.keras.Model(inputs, preprocess_result) # 构造时序模型,加Flatten层适配Dense输出 body = tf.keras.models.Sequential([ tf.keras.layers.Flatten(), tf.keras.layers.Dense(8, activation='relu'), tf.keras.layers.Dense(window_size_y) ]) x_processed = preprocessor(inputs) model_output = body(x_processed) model = tf.keras.Model(inputs, model_output) model.compile(loss='mae', optimizer='adam') # 启动训练,可正常运行 model.fit(dataset.batch(2), epochs=3)
内容的提问来源于stack exchange,提问作者Mykola Zotko
相关产品推荐
相关产品推荐

