You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并异质数据TensorFlow数据集构建窗口数据集报错问题

多数据类型DataFrame构建TensorFlow窗口数据集修复方案

问题背景

需要对包含多种dtype的DataFrame构建窗口化时序数据集,此前同构数据的处理方案在传入字典结构的异构DataFrame时,直接使用flat_map会触发报错:AttributeError: 'dict' object has no attribute 'batch'。
不使用flat_map的初始实现会遇到两类核心问题:

  • 类型错误:窗口生成后直接传入模型触发TypeError: Inputs to a layer should be tensors. Got: <_VariantDataset element_spec=TensorSpec(shape=(), dtype=tf.string, name=None)>,本质是window方法返回的VariantDataset嵌套子数据集结构,没有被转换为模型可接收的张量格式
  • 形状不匹配错误:尝试将嵌套结构转换为张量后,触发StringLookup层报错:ValueError: Exception encountered when calling layer "string_lookup" (type StringLookup). When output_mode is not 'int', maximum supported output rank is 2. Received output_mode one_hot and input shape (None, None), which would result in output rank 3.,本质是窗口化后字符串特征增加了时间步维度,one_hot编码模式默认不支持3维输入

初始问题复现代码

测试数据构造

import pandas as pd
import numpy as np
import tensorflow as tf

# 构造包含字符串、整数两种dtype的测试DataFrame
x = pd.DataFrame({'col1': list('abcdefghij'), 'col2': np.arange(10), 'col3': np.arange(10)})
y = np.arange(10)

初始窗口数据集构建

window_size_x = 3
window_size_y = 2
shift_size = 1

x_slice = x[:-window_size_y]
y_slice = y[window_size_x:]

ds_x = tf.data.Dataset.from_tensor_slices(dict(x_slice)).window(window_size_x, shift=shift_size, drop_remainder=True)
ds_y = tf.data.Dataset.from_tensor_slices(y_slice).window(window_size_y, shift=shift_size, drop_remainder=True)
dataset = tf.data.Dataset.zip((ds_x, ds_y))
dataset = dataset.batch(1)

# 打印单条样本,可见输出为嵌套VariantDataset结构而非张量
for i, j in dataset.take(1):
  print(i, j)

初始预处理器与模型代码

# 构造多输入预处理器
inputs = {'col1': tf.keras.Input(shape=(), name='col1', dtype=tf.string),
          'col2': tf.keras.Input(shape=(), name='col2', dtype=tf.float32),
          'col3': tf.keras.Input(shape=(), name='col3', dtype=tf.float32)}

vocab = sorted(set(x_slice['col1']))
lookup = tf.keras.layers.StringLookup(vocabulary=vocab, output_mode='one_hot')
lookup = lookup(inputs['col1'][:tf.newaxis])

numeric = tf.stack([tf.cast(inputs[i], dtype=tf.float32) for i in ['col2', 'col3']], axis=-1)
preprocess_result = tf.concat([lookup, numeric], axis=-1)

preprocessor = tf.keras.Model(inputs, preprocess_result)
# 直接传入全量切片数据测试,预处理器可正常输出
preprocessor(dict(x_slice))

# 构造端到端模型并尝试训练
body = tf.keras.models.Sequential([tf.keras.layers.Dense(8),
                                   tf.keras.layers.Dense(window_size_y)])
x_processed = preprocessor(inputs)
model_output = body(x_processed)
model = tf.keras.Model(inputs, model_output)

model.compile(loss='mae', optimizer='adam')
# 执行训练触发前述报错
model.fit(dataset)

修复步骤

1. 修复VariantDataset无法转张量的问题

此前直接使用flat_map触发属性错误的原因,是flat_map要求返回的Dataset结构直接支持batch操作,传入字典时无法自动递归处理内部的VariantDataset,必须手动逐列对窗口子数据集做batch转换,将每个窗口聚合为固定形状的张量。
修复后的数据集构建代码:

window_size_x = 3
window_size_y = 2
shift_size = 1

x_slice = x[:-window_size_y]
y_slice = y[window_size_x:]

ds_x = tf.data.Dataset.from_tensor_slices(dict(x_slice)).window(window_size_x, shift=shift_size, drop_remainder=True)
# 逐列处理字典内的特征,将每个窗口的子Dataset batch为形状(window_size_x,)的张量
ds_x = ds_x.map(lambda col_dict: {col_name: window_ds.batch(window_size_x) for col_name, window_ds in col_dict.items()})
ds_y = tf.data.Dataset.from_tensor_slices(y_slice).window(window_size_y, shift=shift_size, drop_remainder=True)
# 标签窗口同步batch为张量
ds_y = ds_y.map(lambda window_ds: window_ds.batch(window_size_y))

dataset = tf.data.Dataset.zip((ds_x, ds_y))
# 验证单条样本输出,此时返回标准张量
for x_batch, y_batch in dataset.take(1):
    print({k: v.shape for k, v in x_batch.items()}, y_batch.shape)

运行验证输出:

{'col1': TensorShape([3]), 'col2': TensorShape([3]), 'col3': TensorShape([3])} TensorShape([2])

2. 修复StringLookup层3维输入报错

窗口化后单样本特征形状为(window_size_x,),比预处理器初始测试的单步输入多了时间步维度,需要调整输入层形状适配窗口结构,同时调整字符串编码逻辑:直接对形状为(batch_size, window_size_x)的字符串输入做one_hot编码,输出维度自动对齐为(batch_size, window_size_x, vocab_size+1),和扩维后的数值特征拼接即可。
修复后的预处理器与模型代码:

# 调整输入层形状,适配窗口化后(窗口长度,)的单样本特征维度
inputs = {'col1': tf.keras.Input(shape=(window_size_x,), name='col1', dtype=tf.string),
          'col2': tf.keras.Input(shape=(window_size_x,), name='col2', dtype=tf.float32),
          'col3': tf.keras.Input(shape=(window_size_x,), name='col3', dtype=tf.float32)}

vocab = sorted(set(x_slice['col1']))
lookup_layer = tf.keras.layers.StringLookup(vocabulary=vocab, output_mode='one_hot')
# 对字符串特征做one_hot编码,输出形状为(窗口长度, 词表大小+1)
col1_encoded = lookup_layer(inputs['col1'])

# 处理数值特征,对齐维度后和字符串编码结果拼接
col2_cast = tf.cast(tf.expand_dims(inputs['col2'], axis=-1), tf.float32)
col3_cast = tf.cast(tf.expand_dims(inputs['col3'], axis=-1), tf.float32)
preprocess_result = tf.concat([col1_encoded, col2_cast, col3_cast], axis=-1)

preprocessor = tf.keras.Model(inputs, preprocess_result)

# 构造时序模型,加Flatten层适配Dense输出
body = tf.keras.models.Sequential([
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(8, activation='relu'),
    tf.keras.layers.Dense(window_size_y)
])
x_processed = preprocessor(inputs)
model_output = body(x_processed)
model = tf.keras.Model(inputs, model_output)

model.compile(loss='mae', optimizer='adam')
# 启动训练,可正常运行
model.fit(dataset.batch(2), epochs=3)

内容的提问来源于stack exchange,提问作者Mykola Zotko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 12:48:15