You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何带partition_by参数将Polars DataFrame写入S3?

问题描述

当partition_by=None时,可通过以下代码将Polars DataFrame写入S3:

import os
import s3fs
import polars as pl

df = pl.DataFrame({'a': [1, 2, 3, 4], 'b': ['left', 'left', 'right', 'right']})

os.environ['AWS_CA_BUNDLE'] = r'C:\my\path\to\cert.pem'

fs = s3fs.S3FileSystem(
    key='my-key',
    secret='my-secret',
    endpoint_url='https://my-endpoint-url.net'
)

with fs.open('s3://my-bucket-name/my/path/to/file', mode='wb') as f:
    df.write_parquet(f)

但设置partition_by='b'并使用相同的文件对象写入方式时:

import os
import s3fs
import polars as pl

df = pl.DataFrame({'a': [1, 2, 3, 4], 'b': ['left', 'left', 'right', 'right']})

os.environ['AWS_CA_BUNDLE'] = r'C:\my\path\to\cert.pem'

fs = s3fs.S3FileSystem(
    key='my-key',
    secret='my-secret',
    endpoint_url='https://my-endpoint-url.net'
)

with fs.open('s3://my-bucket-name/my/path/to/file', mode='wb') as f:
    df.write_parquet(f, partition_by='b')

会触发错误:

TypeError: 'S3File' object cannot be converted to 'PyString'

环境信息:

  • Python 3.9.7
  • Polars 1.26.0
  • s3fs 2025.3.2
  • aiobotocore 2.17.0
  • botocore 1.35.93
  • boto3 1.35.93
解决方案

原因说明

使用partition_by参数时,Polars会按指定列拆分数据,生成多个分区子目录(如b=left/、b=right/)及对应的Parquet文件,因此无法写入单个S3File对象,必须指定一个目录路径,由Polars自动处理分区目录和文件的创建。

方法一:直接通过S3路径写入(推荐)

无需手动创建s3fs对象,通过storage_options传递AWS配置参数:

import os
import polars as pl

df = pl.DataFrame({'a': [1, 2, 3, 4], 'b': ['left', 'left', 'right', 'right']})

# 设置CA证书环境变量
os.environ['AWS_CA_BUNDLE'] = r'C:\my\path\to\cert.pem'

# 直接写入S3目录,自动生成分区
df.write_parquet(
    path='s3://my-bucket-name/my/path/to/partitioned_dir',
    partition_by='b',
    storage_options={
        'key': 'my-key',
        'secret': 'my-secret',
        'endpoint_url': 'https://my-endpoint-url.net'
    }
)

方法二:使用已创建的s3fs对象

如果需要复用已配置的s3fs文件系统,可通过storage_options传入该对象:

import os
import s3fs
import polars as pl

df = pl.DataFrame({'a': [1, 2, 3, 4], 'b': ['left', 'left', 'right', 'right']})

os.environ['AWS_CA_BUNDLE'] = r'C:\my\path\to\cert.pem'

fs = s3fs.S3FileSystem(
    key='my-key',
    secret='my-secret',
    endpoint_url='https://my-endpoint-url.net'
)

# 传入已创建的s3fs对象写入分区
df.write_parquet(
    path='s3://my-bucket-name/my/path/to/partitioned_dir',
    partition_by='b',
    storage_options={'fs': fs}
)

执行后,S3目标目录下会生成如下结构:

partitioned_dir/
├── b=left/
│   └── data.parquet
└── b=right/
    └── data.parquet

内容的提问来源于stack exchange,提问作者chengcj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 09:37:36