You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为PyArrow数据集正确指定JSON的block_size参数?

问题:将换行分隔的JSON文件读取为PyArrow数据集时,如何正确指定JSON的block_size?

背景

需要将换行分隔的JSON文件转换为Parquet格式。某类文件可通过以下代码正常处理:

# ... 导入语句、schema定义

dataset = ds.dataset(
    'data/landing/type_one',
    format='json',
    schema=schema_type_one
)
ds.write_dataset(
    dataset, 
    'data/intermediate/type_one', 
    format='parquet',
    min_rows_per_group=3*512**2,
    max_rows_per_file=5*10**6
)

但处理另一类包含超长行的文件时,触发错误:pyarrow.lib.ArrowInvalid: straddling object straddles two block boundaries (try to increase block size?),经排查需要增大block_size。

目前通过导入私有APIpyarrow._dataset.JsonFileFormat实现了功能,但希望使用更规范的公开API方式:

import pyarrow.dataset as ds
from pyarrow._dataset import JsonFileFormat
from pyarrow.json import ReadOptions

# ... schema定义

fileformat = JsonFileFormat(
    read_options=ReadOptions(block_size=5*2**20)
)

dataset = ds.dataset(
    'data/landing/type_two',
    format=fileformat,
    schema=schema_type_two
)

ds.write_dataset(
    dataset, 
    'data/intermediate/type_two', 
    format='parquet',
    min_rows_per_group=3*512**2,
    max_rows_per_file=5*10**6
)

解答

方法一:使用公开的ds.JsonFileFormat(推荐,PyArrow 14.0.0+)

从PyArrow 14.0.0版本开始,pyarrow.dataset正式提供了公开的JsonFileFormat类,无需依赖私有API。直接通过该类实例化并传入read_options即可:

import pyarrow.dataset as ds
from pyarrow.json import ReadOptions

# ... schema定义

# 创建公开API的JsonFileFormat实例,设置block_size为5MB
file_format = ds.JsonFileFormat(
    read_options=ReadOptions(block_size=5*2**20)
)

dataset = ds.dataset(
    'data/landing/type_two',
    format=file_format,
    schema=schema_type_two
)

ds.write_dataset(
    dataset, 
    'data/intermediate/type_two', 
    format='parquet',
    min_rows_per_group=3*512**2,
    max_rows_per_file=5*10**6
)

方法二:通过format_options参数传递配置(兼容旧版本)

如果你的PyArrow版本低于14.0.0,可直接在ds.dataset中通过format_options参数传递JSON读取配置,无需显式创建格式对象:

import pyarrow.dataset as ds
from pyarrow.json import ReadOptions

# ... schema定义

dataset = ds.dataset(
    'data/landing/type_two',
    format='json',
    schema=schema_type_two,
    # 直接通过字典传递read_options
    format_options={"read_options": ReadOptions(block_size=5*2**20)}
)

ds.write_dataset(
    dataset, 
    'data/intermediate/type_two', 
    format='parquet',
    min_rows_per_group=3*512**2,
    max_rows_per_file=5*10**6
)

两种方式均使用公开API,避免了依赖私有模块带来的兼容性风险。


内容的提问来源于stack exchange,提问作者teejay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 12:50:04