You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按行读取含嵌套结构的Parquet文件而非按数组元素?

解决TensorFlow-IO读取含嵌套列表的Parquet文件时保留原行结构的问题

问题核心是tfio.IODataset.from_parquet默认会将Parquet中的重复(数组/列表)字段展开,把每个数组元素拆成单独行并重复其他列值,导致出现笛卡尔积。要保留原行结构,需要显式指定schema标记列表字段为重复类型。

解决方案1:手动定义Schema

如果已知Parquet文件的字段类型,直接手动定义schema,将列表字段设为VarLenFeature:

import tensorflow as tf
import tensorflow_io as tfio
from typing import Optional, List

def load_single_parquet_file(source_file_path: str, relevant_features_list: Optional[List[str]]) -> tf.data.Dataset:
    # 根据实际字段类型定义schema,列表字段用VarLenFeature
    schema = {
        "timestamp": tf.io.FixedLenFeature([], tf.int64),  # 可根据实际timestamp格式调整类型
        "device_type": tf.io.VarLenFeature(tf.string),     # 字符串列表列
        "previous_scores": tf.io.VarLenFeature(tf.string)  # 字符串列表列
    }

    # 过滤只保留指定列
    if relevant_features_list is not None:
        schema = {k: v for k, v in schema.items() if k in relevant_features_list}

    # 加载数据集时传入schema
    single_file_dataset = tfio.IODataset.from_parquet(
        source_file_path,
        columns=relevant_features_list,
        schema=schema
    )

    # 将稀疏张量转换为密集一维张量,还原列表结构
    def process_row(row):
        processed_row = {}
        for key, value in row.items():
            if isinstance(value, tf.sparse.SparseTensor):
                processed_row[key] = tf.sparse.to_dense(value)
            else:
                processed_row[key] = value
        return processed_row

    return single_file_dataset.map(process_row)

解决方案2:自动读取并转换Schema

如果不清楚字段类型,可以先读取Parquet的原始schema,自动识别重复字段并转换:

import tensorflow as tf
import tensorflow_io as tfio
from typing import Optional, List

def load_single_parquet_file(source_file_path: str, relevant_features_list: Optional[List[str]]) -> tf.data.Dataset:
    # 读取Parquet文件的原始schema
    raw_schema = tfio.parquet.read_schema(source_file_path)

    # 转换为TensorFlow兼容的schema,标记重复字段为VarLenFeature
    schema = {}
    for field in raw_schema:
        if relevant_features_list is not None and field.name not in relevant_features_list:
            continue
        # 判断是否为重复(列表)类型
        if field.repetition_type == tfio.parquet.ParquetRepetitionType.REPEATED:
            schema[field.name] = tf.io.VarLenFeature(field.dtype)
        else:
            schema[field.name] = tf.io.FixedLenFeature([], field.dtype)

    # 加载数据集
    single_file_dataset = tfio.IODataset.from_parquet(
        source_file_path,
        columns=relevant_features_list,
        schema=schema
    )

    # 稀疏转密集,保留列表结构
    def process_row(row):
        return {
            k: tf.sparse.to_dense(v) if isinstance(v, tf.sparse.SparseTensor) else v 
            for k, v in row.items()
        }

    return single_file_dataset.map(process_row)

关键说明

  • 通过指定schema参数,告诉TensorFlow-IO要保留重复字段的嵌套结构,而非展开为笛卡尔积行。
  • VarLenFeature会将列表字段读取为SparseTensor,最后通过tf.sparse.to_dense转换为常规的一维张量,还原原行的数组结构。
  • 该方案适配tensorflow-io 0.36与tensorflow 2.15的版本组合。

内容的提问来源于stack exchange,提问作者André Claudino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 13:10:13