如何按指定维度拼接TensorFlow两个时序数据集的对应批次
你可以通过TensorFlow数据集的zip配对+map拼接逻辑实现需求,完整实现代码如下:
import tensorflow as tf import numpy as np simple_features = np.array([ [1, 1, 1], [2, 2, 2], [3, 3, 3], [4, 4, 4], [5, 5, 5], ]) simple_labels = np.array([ [-1, -1], [-2, -2], [-3, -3], [-4, -4], [-5, -5], ]) simple_features1 = np.array([ [1, 4, 1], [2, 2, 2], [3, 3, 3], [6, 4, 4], [5, 4, 5], ]) simple_labels1 = np.array([ [8, -7], [-2, -2], [-3, 7], [-4, 9], [-5, -5], ]) def print_dataset(ds): for inputs, targets in ds: print("---Batch---") print("Feature:", inputs.numpy()) print("Label:", targets.numpy()) print("") ds1 = tf.keras.preprocessing.timeseries_dataset_from_array(simple_features, simple_labels, sequence_length=4, batch_size=1) ds2 = tf.keras.preprocessing.timeseries_dataset_from_array(simple_features1, simple_labels1, sequence_length=4, batch_size=1) # 合并数据集核心逻辑 combined_ds = tf.data.Dataset.zip((ds1, ds2)).map( lambda x, y: (tf.concat([x[0], y[0]], axis=-1), tf.concat([x[1], y[1]], axis=-1)) ) # 打印验证合并结果 print_dataset(combined_ds)
运行后输出结果和你预期完全一致,第一个批次输出如下:
---Batch--- Feature: [[[1 1 1 1 4 1] [2 2 2 2 2 2] [3 3 3 3 3 3] [4 4 4 6 4 4]]] Label: [[-1 -1 8 -7]]
其中axis=-1指定在最后一维拼接,刚好满足特征、标签横向合并的需求,zip方法会严格按照两个数据集的批次顺序配对,只要两个数据集批次数量一致就可以稳定运行。
内容的提问来源于stack exchange,提问作者freak11
相关产品推荐
相关产品推荐

