You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用boto3无凭证访问Common Crawl S3桶遇403/访问拒绝错误

问题描述

尝试无需凭证,使用boto3从Common Crawl获取免费可访问的文件,但运行代码时出现403 Forbidden错误。

报错的下载代码

def main():
    s3 = boto3.client("s3", config=Config(signature_version=UNSIGNED))

    s3.download_file("commoncrawl", "crawl-data/CC-MAIN-2022-05/wet.paths.gz", "wet.paths.gz")

下载时的错误信息

botocore.exceptions.ClientError: An error occurred (403) when calling the HeadObject operation: Forbidden

调试时的列表操作及错误

尝试列出Common Crawl桶内内容也失败,两种方式均返回访问拒绝:

方式1:使用client列出对象

s3 = boto3.client("s3", config=Config(signature_version=UNSIGNED))
s3.list_objects(Bucket="commoncrawl")

方式2:使用resource遍历对象

s3 = boto3.resource("s3", config=Config(signature_version=UNSIGNED))
crawl = s3.Bucket('commoncrawl')
for obj in crawl.objects.all():
    print(obj)

列表操作的错误信息

botocore.exceptions.ClientError: An error occurred (AccessDenied) when calling the ListObjects operation: Access Denied

解决方法
  • Common Crawl的S3桶不支持匿名列出对象,桶策略直接禁止匿名用户的ListObjects权限,所以所有遍历桶内容的操作都会返回AccessDenied,这是正常限制。
  • 下载报错是因为download_file默认会调用HeadObject检查文件元数据,而该桶的匿名权限可能未开放此操作。可以改用get_object直接获取文件内容后写入本地,绕开HeadObject检查:
import boto3
from botocore.config import Config

def main():
    s3 = boto3.client("s3", config=Config(signature_version=UNSIGNED))
    # 直接获取对象内容
    response = s3.get_object(Bucket="commoncrawl", Key="crawl-data/CC-MAIN-2022-05/wet.paths.gz")
    
    # 将内容写入本地文件
    with open("wet.paths.gz", "wb") as f:
        f.write(response["Body"].read())

if __name__ == "__main__":
    main()
  • 额外注意:必须明确知道目标文件的完整路径才能匿名访问,无法通过遍历桶来查找文件;若仍报错,检查网络环境是否有代理/防火墙限制S3访问,或确认文件路径是否正确。

内容的提问来源于stack exchange,提问作者Nick H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 16:04:59