如何按行读取含嵌套结构的Parquet文件而非按数组元素?
解决TensorFlow-IO读取含嵌套列表的Parquet文件时保留原行结构的问题
问题核心是tfio.IODataset.from_parquet默认会将Parquet中的重复(数组/列表)字段展开,把每个数组元素拆成单独行并重复其他列值,导致出现笛卡尔积。要保留原行结构,需要显式指定schema标记列表字段为重复类型。
解决方案1:手动定义Schema
如果已知Parquet文件的字段类型,直接手动定义schema,将列表字段设为VarLenFeature:
import tensorflow as tf import tensorflow_io as tfio from typing import Optional, List def load_single_parquet_file(source_file_path: str, relevant_features_list: Optional[List[str]]) -> tf.data.Dataset: # 根据实际字段类型定义schema,列表字段用VarLenFeature schema = { "timestamp": tf.io.FixedLenFeature([], tf.int64), # 可根据实际timestamp格式调整类型 "device_type": tf.io.VarLenFeature(tf.string), # 字符串列表列 "previous_scores": tf.io.VarLenFeature(tf.string) # 字符串列表列 } # 过滤只保留指定列 if relevant_features_list is not None: schema = {k: v for k, v in schema.items() if k in relevant_features_list} # 加载数据集时传入schema single_file_dataset = tfio.IODataset.from_parquet( source_file_path, columns=relevant_features_list, schema=schema ) # 将稀疏张量转换为密集一维张量,还原列表结构 def process_row(row): processed_row = {} for key, value in row.items(): if isinstance(value, tf.sparse.SparseTensor): processed_row[key] = tf.sparse.to_dense(value) else: processed_row[key] = value return processed_row return single_file_dataset.map(process_row)
解决方案2:自动读取并转换Schema
如果不清楚字段类型,可以先读取Parquet的原始schema,自动识别重复字段并转换:
import tensorflow as tf import tensorflow_io as tfio from typing import Optional, List def load_single_parquet_file(source_file_path: str, relevant_features_list: Optional[List[str]]) -> tf.data.Dataset: # 读取Parquet文件的原始schema raw_schema = tfio.parquet.read_schema(source_file_path) # 转换为TensorFlow兼容的schema,标记重复字段为VarLenFeature schema = {} for field in raw_schema: if relevant_features_list is not None and field.name not in relevant_features_list: continue # 判断是否为重复(列表)类型 if field.repetition_type == tfio.parquet.ParquetRepetitionType.REPEATED: schema[field.name] = tf.io.VarLenFeature(field.dtype) else: schema[field.name] = tf.io.FixedLenFeature([], field.dtype) # 加载数据集 single_file_dataset = tfio.IODataset.from_parquet( source_file_path, columns=relevant_features_list, schema=schema ) # 稀疏转密集,保留列表结构 def process_row(row): return { k: tf.sparse.to_dense(v) if isinstance(v, tf.sparse.SparseTensor) else v for k, v in row.items() } return single_file_dataset.map(process_row)
关键说明
- 通过指定
schema参数,告诉TensorFlow-IO要保留重复字段的嵌套结构,而非展开为笛卡尔积行。 VarLenFeature会将列表字段读取为SparseTensor,最后通过tf.sparse.to_dense转换为常规的一维张量,还原原行的数组结构。- 该方案适配tensorflow-io 0.36与tensorflow 2.15的版本组合。
内容的提问来源于stack exchange,提问作者André Claudino
相关产品推荐
相关产品推荐

