Scrapy自定义Wasabi S3管道遇InvalidAccessKeyId错误求助
问题:Scrapy自定义Wasabi S3管道上传成功,但仍报InvalidAccessKeyId错误
- 自定义Scrapy管道可正常将JSON文件上传至Wasabi S3,但日志中持续出现错误:
botocore.exceptions.ClientError: An error occurred (InvalidAccessKeyId) when calling the PutObject operation: The AWS Access Key Id you provided does not exist in our records. - 已验证访问密钥有效,通过终端和单独使用boto3上传文件无异常,且账号配置了
WasabiFullAccess和AmazonS3FullAccess权限
自定义管道代码
from scrapy import signals from scrapy.exporters import JsonItemExporter import boto3 class JsonWriterPipeline(object): def __init__(self): self.items_list = [] @classmethod def from_crawler(cls, crawler): pipeline = cls() crawler.signals.connect(pipeline.spider_opened, signals.spider_opened) crawler.signals.connect(pipeline.spider_closed, signals.spider_closed) return pipeline def spider_opened(self, spider): self.file = open("%s_items.json" % spider.name, "wb") self.exporter = JsonItemExporter(self.file) self.exporter.encoding = "utf-8" self.exporter.start_exporting() def process_item(self, item, spider): self.exporter.export_item(item) return item def spider_closed(self, spider): self.exporter.finish_exporting() self.file.close() s3 = boto3.resource( "s3", endpoint_url="https://s3.ap-northeast-1.wasabisys.com", aws_access_key_id="ACCESS_KEY_ID", aws_secret_access_key="SECRET_ACCESS_KEY", ) boto_test_bucket = s3.Bucket(bucket_name) boto_test_bucket.upload_file("%s_items.json" % spider.name, f"{spider.name}")
错误回溯日志
2022-10-31 03:50:18 [scrapy.extensions.feedexport] ERROR: Error storing json feed (120 items) in: s3://brownfashions/brown_fashion_json/brown_fashion_json_2022-10-30T22-50-00.json Traceback (most recent call last): File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/threadpool.py", line 244, in inContext result = inContext.theWork() # type: ignore[attr-defined] File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/threadpool.py", line 260, in <lambda> inContext.theWork = lambda: context.call( # type: ignore[attr-defined] File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/context.py", line 117, in callWithContext return self.currentContext().callWithContext(ctx, func, *args, **kw) File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/twisted/python/context.py", line 82, in callWithContext return func(*args, **kw) File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/scrapy/extensions/feedexport.py", line 196, in _store_in_thread self.s3_client.put_object( File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/botocore/client.py", line 507, in _api_call return self._make_api_call(operation_name, kwargs) File "/home/zerox/BrownFashions/venv/lib/python3.9/site-packages/botocore/client.py", line 943, in _make_api_call raise error_class(parsed_response, operation_name) botocore.exceptions.ClientError: An error occurred (InvalidAccessKeyId) when calling the PutObject operation: The AWS Access Key Id you provided does not exist in our records.
错误原因
报错来源是scrapy.extensions.feedexport,说明错误并非来自你自定义的管道代码,而是Scrapy内置的Feed导出组件。你的自定义管道已经成功完成上传,但Scrapy同时在尝试通过FeedExport功能自动将数据上传到S3,而该功能的密钥/端点配置不正确,导致了InvalidAccessKeyId错误。
解决办法
方案1:禁用FeedExport功能
如果不需要自动导出Feed到S3,直接在settings.py中删除或注释相关配置:
# 找到如下配置并删除/注释 FEEDS = { 's3://brownfashions/brown_fashion_json/%(name)s_json_%(time)s.json': { 'format': 'json', 'encoding': 'utf8', } }
方案2:为FeedExport配置正确的Wasabi参数
若需要保留FeedExport功能,在settings.py中配置适配Wasabi的S3参数:
# FeedExport配置 FEEDS = { 's3://brownfashions/brown_fashion_json/%(name)s_json_%(time)s.json': { 'format': 'json', 'encoding': 'utf8', } } # Wasabi S3客户端配置 AWS_ACCESS_KEY_ID = '你的Wasabi访问密钥ID' AWS_SECRET_ACCESS_KEY = '你的Wasabi秘密访问密钥' AWS_ENDPOINT_URL = 'https://s3.ap-northeast-1.wasabisys.com'
优化建议
自定义管道中硬编码密钥的做法不安全,建议改为从settings.py读取:
# 在spider_closed方法中修改 from scrapy.utils.project import get_project_settings def spider_closed(self, spider): self.exporter.finish_exporting() self.file.close() settings = get_project_settings() s3 = boto3.resource( "s3", endpoint_url=settings.get('AWS_ENDPOINT_URL'), aws_access_key_id=settings.get('AWS_ACCESS_KEY_ID'), aws_secret_access_key=settings.get('AWS_SECRET_ACCESS_KEY'), ) boto_test_bucket = s3.Bucket(bucket_name) boto_test_bucket.upload_file("%s_items.json" % spider.name, f"{spider.name}")
内容的提问来源于stack exchange,提问作者X-something
相关产品推荐
相关产品推荐

