You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy自定义Wasabi S3管道遇InvalidAccessKeyId错误求助

问题:Scrapy自定义Wasabi S3管道上传成功,但仍报InvalidAccessKeyId错误
  • 自定义Scrapy管道可正常将JSON文件上传至Wasabi S3,但日志中持续出现错误:botocore.exceptions.ClientError: An error occurred (InvalidAccessKeyId) when calling the PutObject operation: The AWS Access Key Id you provided does not exist in our records.
  • 已验证访问密钥有效,通过终端和单独使用boto3上传文件无异常,且账号配置了WasabiFullAccess和AmazonS3FullAccess权限

自定义管道代码

from scrapy import signals
from scrapy.exporters import JsonItemExporter
import boto3


class JsonWriterPipeline(object):
    def __init__(self):
        self.items_list = []

    @classmethod
    def from_crawler(cls, crawler):
        pipeline = cls()
        crawler.signals.connect(pipeline.spider_opened, signals.spider_opened)
        crawler.signals.connect(pipeline.spider_closed, signals.spider_closed)
        return pipeline

    def spider_opened(self, spider):
        self.file = open("%s_items.json" % spider.name, "wb")
        self.exporter = JsonItemExporter(self.file)
        self.exporter.encoding = "utf-8"
        self.exporter.start_exporting()

    def process_item(self, item, spider):
        self.exporter.export_item(item)
        return item

    def spider_closed(self, spider):
        self.exporter.finish_exporting()
        self.file.close()
        s3 = boto3.resource(
            "s3",
            endpoint_url="https://s3.ap-northeast-1.wasabisys.com",
            aws_access_key_id="ACCESS_KEY_ID",
            aws_secret_access_key="SECRET_ACCESS_KEY",
        )
        boto_test_bucket = s3.Bucket(bucket_name)
        boto_test_bucket.upload_file("%s_items.json" % spider.name, f"{spider.name}")

错误回溯日志

2022-10-31 03:50:18 [scrapy.extensions.feedexport] ERROR: Error storing json feed (120 items) in: s3://brownfashions/brown_fashion_json/brown_fashion_json_2022-10-30T22-50-00.json
Traceback (most recent call last):
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/threadpool.py", line 244, in inContext
    result = inContext.theWork()  # type: ignore[attr-defined]
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/threadpool.py", line 260, in <lambda>
    inContext.theWork = lambda: context.call(  # type: ignore[attr-defined]
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/context.py", line 117, in callWithContext
    return self.currentContext().callWithContext(ctx, func, *args, **kw)
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/context.py", line 82, in callWithContext
    return func(*args, **kw)
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/scrapy/extensions/feedexport.py", line 196, in _store_in_thread
    self.s3_client.put_object(
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/botocore/client.py", line 507, in _api_call
    return self._make_api_call(operation_name, kwargs)
  File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/botocore/client.py", line 943, in _make_api_call
    raise error_class(parsed_response, operation_name)
botocore.exceptions.ClientError: An error occurred (InvalidAccessKeyId) when calling the PutObject operation: The AWS Access Key Id you provided does not exist in our records.

错误原因

报错来源是scrapy.extensions.feedexport,说明错误并非来自你自定义的管道代码,而是Scrapy内置的Feed导出组件。你的自定义管道已经成功完成上传,但Scrapy同时在尝试通过FeedExport功能自动将数据上传到S3,而该功能的密钥/端点配置不正确,导致了InvalidAccessKeyId错误。

解决办法

方案1:禁用FeedExport功能

如果不需要自动导出Feed到S3,直接在settings.py中删除或注释相关配置:

# 找到如下配置并删除/注释
FEEDS = {
    's3://brownfashions/brown_fashion_json/%(name)s_json_%(time)s.json': {
        'format': 'json',
        'encoding': 'utf8',
    }
}

方案2:为FeedExport配置正确的Wasabi参数

若需要保留FeedExport功能,在settings.py中配置适配Wasabi的S3参数:

# FeedExport配置
FEEDS = {
    's3://brownfashions/brown_fashion_json/%(name)s_json_%(time)s.json': {
        'format': 'json',
        'encoding': 'utf8',
    }
}

# Wasabi S3客户端配置
AWS_ACCESS_KEY_ID = '你的Wasabi访问密钥ID'
AWS_SECRET_ACCESS_KEY = '你的Wasabi秘密访问密钥'
AWS_ENDPOINT_URL = 'https://s3.ap-northeast-1.wasabisys.com'

优化建议

自定义管道中硬编码密钥的做法不安全,建议改为从settings.py读取:

# 在spider_closed方法中修改
from scrapy.utils.project import get_project_settings

def spider_closed(self, spider):
    self.exporter.finish_exporting()
    self.file.close()
    settings = get_project_settings()
    s3 = boto3.resource(
        "s3",
        endpoint_url=settings.get('AWS_ENDPOINT_URL'),
        aws_access_key_id=settings.get('AWS_ACCESS_KEY_ID'),
        aws_secret_access_key=settings.get('AWS_SECRET_ACCESS_KEY'),
    )
    boto_test_bucket = s3.Bucket(bucket_name)
    boto_test_bucket.upload_file("%s_items.json" % spider.name, f"{spider.name}")

内容的提问来源于stack exchange,提问作者X-something

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:31:03