You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyarrow的write_dataset函数时,能否指定压缩格式?

控制PyArrow分区写入时的Parquet压缩类型

完全支持自定义压缩类型,你只需要给ds.write_dataset添加format_options参数,通过ParquetWriteOptions指定想要的压缩方式,就能替换默认的snappy。

修改后的代码示例:

import numpy.random
import pyarrow as pa
import pyarrow.dataset as ds

data = pa.table(
    {
        "day": numpy.random.randint(1, 31, size=100),
        "month": numpy.random.randint(1, 12, size=100),
        "year": [2000 + x // 10 for x in range(100)],
    }
)

# 指定压缩类型为gzip,可替换为zstd、brotli、lz4等
write_options = pa.parquet.ParquetWriteOptions(compression='gzip')

ds.write_dataset(
    data,
    "./tmp/partitioned",
    format="parquet",
    existing_data_behavior="delete_matching",
    partitioning=ds.partitioning(
        pa.schema(
            [
                ("year", pa.int16()),
            ]
        ),
    ),
    format_options=write_options  # 传入压缩配置
)

你可以根据需求把compression参数换成PyArrow支持的其他压缩算法,比如'zstd'(高效压缩推荐)、'brotli'、'lz4'或者'uncompressed'(不压缩)。

内容的提问来源于stack exchange,提问作者MBr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:15:34