在AWS Glue作业中上传pandas-profiling HTML输出至S3出现空文件问题的代码调整咨询
解决AWS Glue中pandas-profiling HTML上传S3为空的问题
我帮你排查了代码里的核心问题,主要出在内存流的选择和指针处理上,调整以下几点就能搞定空文件的问题:
问题分析
- 错误使用StringIO处理二进制内容:pandas-profiling生成的HTML输出是二进制格式的,
StringIO是专门用于文本流的容器,用它存储二进制内容会导致编码异常,最终写入S3的内容为空,应该换成BytesIO来处理二进制流。 - 未重置文件指针位置:当你向内存流写入内容后,流的指针会停在末尾,直接调用
getvalue()会读取到空内容,需要先将指针移到流的起始位置。 - 冗余的本地文件写入:在Glue作业中,
profile.to_file()写入本地临时文件完全没必要,既浪费资源又没有实际作用,直接删除这一步即可。
修正后的代码
import pandas as pd import boto3 import io from pandas_profiling import ProfileReport from io import BytesIO # 替换StringIO为BytesIO #Pull all file names/keys from S3 s3 = boto3.client('s3') def get_matching_s3_keys(bucket, prefix='', suffix=''): """ Generate the keys in an S3 bucket. :param bucket: Name of the S3 bucket. :param prefix: Only fetch keys that start with this prefix (optional). :param suffix: Only fetch keys that end with this suffix (optional). """ kwargs = {'Bucket': bucket, 'Prefix': prefix} while True: resp = s3.list_objects_v2(**kwargs) for obj in resp['Contents']: key = obj['Key'] if key.endswith(suffix): yield key try: kwargs['ContinuationToken'] = resp['NextContinuationToken'] except KeyError: break #Pull all file paths and append to list tables_list = [] for key in get_matching_s3_keys('mybucketname', 'processed/', '.csv'): print(key) tables_list.append(key) for i in tables_list: obj = s3.get_object(Bucket='mybucketname', Key=i) df = pd.read_csv(obj['Body']) profile = ProfileReport(df, title = 'My Data Profile', html ={"style": {'full_width':True}}, minimal=True) # 移除本地文件写入的冗余代码 # profile.to_file(i.lstrip("processed/").rstrip(".csv")+".html") # 使用BytesIO处理二进制流 buf = BytesIO() profile.to_file(buf, 'html') # 写入二进制流 buf.seek(0) # 将指针移到流的开头 # Upload as bytes s3.put_object( Bucket='mybucketname', Key=i.lstrip("processed/").rstrip(".csv")+".html", Body=buf )
额外提示
- 如果你需要验证内存流中的内容,可以在
buf.seek(0)后添加print(buf.read()),查看是否有正确的HTML内容输出(测试完成后记得注释掉)。 - 在Glue作业中,确保IAM角色拥有S3的读写权限,避免因为权限问题导致上传失败(不过你的情况是空文件,所以权限应该没问题,但还是可以确认下)。
内容的提问来源于stack exchange,提问作者pdangelo4
相关产品推荐
相关产品推荐

