You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas写入Parquet时如何实现按列压缩?

Pandas结合Parquet实现按列压缩的方法

Pandas原生的DataFrame.to_parquet()确实只支持全局压缩设置,没法直接传入按列配置的字典,但可以通过结合PyArrow的底层API实现按列压缩,步骤很直接:

实现步骤

  1. 将Pandas DataFrame转换为PyArrow Table对象
  2. 使用PyArrow的pq.write_table()方法,传入按列定义压缩策略的字典

代码示例

import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq

# 创建示例DataFrame
df = pd.DataFrame({
    'text_column': ['hello', 'world'] * 500,
    'int_column': range(1000),
    'float_column': [3.14, 2.71] * 500
})

# 转换为PyArrow Table
arrow_table = pa.Table.from_pandas(df)

# 定义按列压缩策略:文本列用gzip,整数列用snappy,浮点列不压缩
column_compression = {
    'text_column': 'gzip',
    'int_column': 'snappy',
    'float_column': None
}

# 写入按列压缩的Parquet文件
pq.write_table(arrow_table, 'column_compressed.parquet', compression=column_compression)

补充说明

  • 读取该文件时,不管用PyArrow的pq.read_table()还是Pandas的pd.read_parquet(),都能正常解析,不需要额外配置压缩参数
  • Pandas的to_parquet()即使指定engine='pyarrow',也仅支持传入全局压缩字符串(如'snappy'),无法直接传递按列压缩字典,所以必须手动走PyArrow的写入流程

内容的提问来源于stack exchange,提问作者silgon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 11:10:58