You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用awswrangler将Pandas Profile报告输出至AWS S3

解决方法
  • 借助ProfileReport.to_html()生成HTML内容字符串,再通过awswrangler.s3.write直接将内容写入S3存储桶,无需先写入本地文件,节省IO开销。

完整代码示例

import pandas_profiling
import awswrangler as wr

# 假设已生成目标DataFrame对象df
profile = pandas_profiling.ProfileReport(
    df, title=f"file_name Data Profile Report", minimal=True
)

# 生成HTML内容字符串
html_content = profile.to_html()

# 定义S3目标路径,替换为你的存储桶和路径
s3_target_path = "s3://your-bucket-name/processedData/file_name-profile.html"

# 写入S3
wr.s3.write(
    df=html_content,
    path=s3_target_path,
    mode="w",
    encoding="utf-8"
)

备选字节流写法

如果更倾向于处理字节流,也可以用BytesIO包装后写入:

import pandas_profiling
import awswrangler as wr
from io import BytesIO

# 生成ProfileReport
profile = pandas_profiling.ProfileReport(
    df, title=f"file_name Data Profile Report", minimal=True
)

# 将报告写入字节缓冲区
html_buffer = BytesIO()
profile.to_file(html_buffer)
html_buffer.seek(0)  # 重置缓冲区指针到起始位置

# 写入S3
s3_target_path = "s3://your-bucket-name/processedData/file_name-profile.html"
wr.s3.write(
    df=html_buffer.getvalue(),
    path=s3_target_path,
    mode="wb"
)

注意事项

  • 确保运行代码的EC2实例拥有访问目标S3存储桶的IAM权限(至少需要s3:PutObject权限)
  • 替换代码中的s3://your-bucket-name/...为实际的S3路径

内容的提问来源于stack exchange,提问作者Farooque

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:35:16