如何为PyArrow数据集正确指定JSON的block_size参数?
问题:将换行分隔的JSON文件读取为PyArrow数据集时,如何正确指定JSON的
block_size? 背景
需要将换行分隔的JSON文件转换为Parquet格式。某类文件可通过以下代码正常处理:
# ... 导入语句、schema定义 dataset = ds.dataset( 'data/landing/type_one', format='json', schema=schema_type_one ) ds.write_dataset( dataset, 'data/intermediate/type_one', format='parquet', min_rows_per_group=3*512**2, max_rows_per_file=5*10**6 )
但处理另一类包含超长行的文件时,触发错误:pyarrow.lib.ArrowInvalid: straddling object straddles two block boundaries (try to increase block size?),经排查需要增大block_size。
目前通过导入私有APIpyarrow._dataset.JsonFileFormat实现了功能,但希望使用更规范的公开API方式:
import pyarrow.dataset as ds from pyarrow._dataset import JsonFileFormat from pyarrow.json import ReadOptions # ... schema定义 fileformat = JsonFileFormat( read_options=ReadOptions(block_size=5*2**20) ) dataset = ds.dataset( 'data/landing/type_two', format=fileformat, schema=schema_type_two ) ds.write_dataset( dataset, 'data/intermediate/type_two', format='parquet', min_rows_per_group=3*512**2, max_rows_per_file=5*10**6 )
解答
方法一:使用公开的ds.JsonFileFormat(推荐,PyArrow 14.0.0+)
从PyArrow 14.0.0版本开始,pyarrow.dataset正式提供了公开的JsonFileFormat类,无需依赖私有API。直接通过该类实例化并传入read_options即可:
import pyarrow.dataset as ds from pyarrow.json import ReadOptions # ... schema定义 # 创建公开API的JsonFileFormat实例,设置block_size为5MB file_format = ds.JsonFileFormat( read_options=ReadOptions(block_size=5*2**20) ) dataset = ds.dataset( 'data/landing/type_two', format=file_format, schema=schema_type_two ) ds.write_dataset( dataset, 'data/intermediate/type_two', format='parquet', min_rows_per_group=3*512**2, max_rows_per_file=5*10**6 )
方法二:通过format_options参数传递配置(兼容旧版本)
如果你的PyArrow版本低于14.0.0,可直接在ds.dataset中通过format_options参数传递JSON读取配置,无需显式创建格式对象:
import pyarrow.dataset as ds from pyarrow.json import ReadOptions # ... schema定义 dataset = ds.dataset( 'data/landing/type_two', format='json', schema=schema_type_two, # 直接通过字典传递read_options format_options={"read_options": ReadOptions(block_size=5*2**20)} ) ds.write_dataset( dataset, 'data/intermediate/type_two', format='parquet', min_rows_per_group=3*512**2, max_rows_per_file=5*10**6 )
两种方式均使用公开API,避免了依赖私有模块带来的兼容性风险。
内容的提问来源于stack exchange,提问作者teejay
相关产品推荐
相关产品推荐

