boto3下载S3文件报HeadObject 404但CLI操作正常问题排查
问题场景
- 通过boto3提交Athena查询,配置查询结果输出到S3存储桶
- 业务逻辑需要下载S3中存储的Athena查询结果文件到本地
报错信息
调用HeadObject操作时发生错误(404):Not Found
异常现象
- 目标文件实际存在于S3中,执行
aws s3 cp命令可以正常拷贝下载 - 使用boto3下载文件时报上述404错误,底层调用
head-object操作失败 - 执行如下CLI命令可以正常返回对象元数据:
aws s3api head-object --bucket dsp-smaato-sink-prod --key /athena_query_results/c96bdc09-d545-4ee3-bc66-be3be928e3f2.csv
- 已排查账号权限配置,当前使用账号持有管理员全权限,不存在权限缺失问题
关联问题代码
# snippets def s3_donwload(url, target=None): # s3 = boto3.resource('s3') # client = s3.meta.client client = boto3.client("s3", region_name=constant.AWS_REGION, endpoint_url='https://s3.ap-southeast-1.amazonaws.com') s3_file = urlparse(url) if target: target = os.path.abspath(target) else: target = os.path.abspath(os.path.basename(s3_file.path)) logger.info(f"download {url} to {target}...") client.download_file(s3_file.netloc, s3_file.path, target) logger.info(f"download {url} to {target} done!")
根因分析
问题出在S3对象Key的传参格式差异上:
- 用
urlparse解析标准S3路径(格式为s3://bucket-name/path/to/object)时,解析得到的path属性会自带开头的/,示例路径解析后得到的是/athena_query_results/c96bdc09-d545-4ee3-bc66-be3be928e3f2.csv - AWS CLI内置了参数预处理逻辑,会自动剔除传入Key开头的多余斜杠,所以带前导斜杠的CLI命令可以正常执行
- boto3没有做这个自动兼容处理,会把开头带
/的字符串直接作为完整Key去S3中查找对象,而S3中实际存储的对象Key是不带开头斜杠的athena_query_results/c96bdc09-d545-4ee3-bc66-be3be928e3f2.csv,匹配不到对象自然返回404。
修复方案
在传入Key参数前,手动剔除路径开头的前导斜杠即可,修改后的代码逻辑如下:
def s3_donwload(url, target=None): client = boto3.client("s3", region_name=constant.AWS_REGION, endpoint_url='https://s3.ap-southeast-1.amazonaws.com') s3_file = urlparse(url) # 移除S3路径开头的多余前导斜杠 s3_key = s3_file.path.lstrip('/') if target: target = os.path.abspath(target) else: target = os.path.abspath(os.path.basename(s3_key)) logger.info(f"download {url} to {target}...") client.download_file(s3_file.netloc, s3_key, target) logger.info(f"download {url} to {target} done!")
内容的提问来源于stack exchange,提问作者Evans Y.
相关产品推荐
相关产品推荐

