You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在AWS Glue作业中上传pandas-profiling HTML输出至S3出现空文件问题的代码调整咨询

解决AWS Glue中pandas-profiling HTML上传S3为空的问题

我帮你排查了代码里的核心问题,主要出在内存流的选择和指针处理上,调整以下几点就能搞定空文件的问题:

问题分析

  1. 错误使用StringIO处理二进制内容:pandas-profiling生成的HTML输出是二进制格式的,StringIO是专门用于文本流的容器,用它存储二进制内容会导致编码异常,最终写入S3的内容为空,应该换成BytesIO来处理二进制流。
  2. 未重置文件指针位置:当你向内存流写入内容后,流的指针会停在末尾,直接调用getvalue()会读取到空内容,需要先将指针移到流的起始位置。
  3. 冗余的本地文件写入:在Glue作业中,profile.to_file()写入本地临时文件完全没必要,既浪费资源又没有实际作用,直接删除这一步即可。

修正后的代码

import pandas as pd
import boto3
import io
from pandas_profiling import ProfileReport
from io import BytesIO  # 替换StringIO为BytesIO

#Pull all file names/keys from S3
s3 = boto3.client('s3')

def get_matching_s3_keys(bucket, prefix='', suffix=''):
    """
    Generate the keys in an S3 bucket.
    :param bucket: Name of the S3 bucket.
    :param prefix: Only fetch keys that start with this prefix (optional).
    :param suffix: Only fetch keys that end with this suffix (optional).
    """
    kwargs = {'Bucket': bucket, 'Prefix': prefix}
    while True:
        resp = s3.list_objects_v2(**kwargs)
        for obj in resp['Contents']:
            key = obj['Key']
            if key.endswith(suffix):
                yield key
        try:
            kwargs['ContinuationToken'] = resp['NextContinuationToken']
        except KeyError:
            break

#Pull all file paths and append to list
tables_list = []
for key in get_matching_s3_keys('mybucketname', 'processed/', '.csv'):
    print(key)
    tables_list.append(key)

for i in tables_list:
    obj = s3.get_object(Bucket='mybucketname', Key=i)
    df = pd.read_csv(obj['Body'])
    profile = ProfileReport(df, title = 'My Data Profile', html ={"style": {'full_width':True}}, minimal=True)
    
    # 移除本地文件写入的冗余代码
    # profile.to_file(i.lstrip("processed/").rstrip(".csv")+".html")
    
    # 使用BytesIO处理二进制流
    buf = BytesIO()
    profile.to_file(buf, 'html')  # 写入二进制流
    buf.seek(0)  # 将指针移到流的开头
    
    # Upload as bytes
    s3.put_object(
        Bucket='mybucketname',
        Key=i.lstrip("processed/").rstrip(".csv")+".html",
        Body=buf
    )

额外提示

  • 如果你需要验证内存流中的内容,可以在buf.seek(0)后添加print(buf.read()),查看是否有正确的HTML内容输出(测试完成后记得注释掉)。
  • 在Glue作业中,确保IAM角色拥有S3的读写权限,避免因为权限问题导致上传失败(不过你的情况是空文件,所以权限应该没问题,但还是可以确认下)。

内容的提问来源于stack exchange,提问作者pdangelo4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:32:42