如何在Scrapy的FSFilesStore的stat_file方法中实现非阻塞HTTP请求?
问题与解决方案
问题背景
- 需求:通过HEAD请求获取远程文件的大小、最后修改时间,判断是否需要更新文件,避免下载大体积视频文件
- 遇到的问题:扩展
FSFilesStore重写stat_file方法时,使用async/await调用treq出现连接丢失错误,推测该方法不支持async语法 - 痛点:现有方案依赖爬虫实现,过于hack,希望在
stat_file中用非阻塞方式实现HTTP调用,不侵入爬虫逻辑
代码示例(原错误实现)
import treq from scrapy.pipelines.files import FilesPipeline from scrapy.pipelines.files import FSFilesStore class FSFilesStoreDBCacheExtended(FSFilesStore): async def stat_file(self, path, info): try: head_resp = await treq.head('https://google.com') except Exception as e: print(e) # ConnectError here raise e size = int(head_resp.headers['content-length']) mod_timestamp = head_resp.headers['last-modified'] # ... do some stuff return super().stat_file(path, info) class ImageDownload(FilesPipeline): def __init__(self, store_uri, download_func=None, settings=None): self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended super().__init__(store_uri, download_func, settings)
错误信息(中文翻译)
[<twisted.python.failure.Failure twisted.internet.error.ConnectionLost: 连接已非正常断开: 连接丢失。>]
解决方案
Scrapy基于Twisted异步框架,stat_file方法需要返回Deferred对象而非使用async/await(直接用async def会导致事件循环不兼容)。以下提供两种可行方案:
方案1:使用treq(基于Twisted)实现非阻塞HEAD请求
利用treq返回Deferred的特性,通过回调链处理响应和异常:
import treq from scrapy.pipelines.files import FilesPipeline, FSFilesStore class FSFilesStoreDBCacheExtended(FSFilesStore): def stat_file(self, path, info): # 从info中获取远程文件URL(FilesPipeline会传递该元数据) remote_url = info.get('url') if not remote_url: return super().stat_file(path, info) def handle_head_response(response): # 解析HEAD响应头 size = int(response.headers.get('content-length', 0)) mod_timestamp = response.headers.get('last-modified') # 此处添加你的更新判断逻辑: # 1. 对比本地文件与远程文件的大小/修改时间 # 2. 若无需更新,可直接返回自定义状态跳过下载 # 执行父类默认逻辑(或根据判断结果返回对应状态) return super(FSFilesStoreDBCacheExtended, self).stat_file(path, info) def handle_failure(failure): # 请求失败时回退到默认逻辑 print(f"HEAD请求失败: {failure.getErrorMessage()}") return super(FSFilesStoreDBCacheExtended, self).stat_file(path, info) # 发送HEAD请求并返回Deferred d = treq.head(remote_url) d.addCallback(handle_head_response) d.addErrback(handle_failure) return d class ImageDownload(FilesPipeline): def __init__(self, store_uri, download_func=None, settings=None): self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended super().__init__(store_uri, download_func, settings)
方案2:复用Scrapy下载器发送HEAD请求(推荐)
直接使用Scrapy内置的下载器,可复用爬虫的UA、代理、重试策略等配置:
from scrapy.pipelines.files import FilesPipeline, FSFilesStore from scrapy.http import Request class FSFilesStoreDBCacheExtended(FSFilesStore): def stat_file(self, path, info): remote_url = info.get('url') if not remote_url: return super().stat_file(path, info) def handle_download_response(response): size = int(response.headers.get('content-length', 0)) mod_timestamp = response.headers.get('last-modified') # 你的更新判断逻辑 return super().stat_file(path, info) def handle_failure(failure): print(f"HEAD请求失败: {failure.getErrorMessage()}") return super().stat_file(path, info) # 构造HEAD请求,通过Scrapy下载器发送 head_req = Request(remote_url, method='HEAD') d = info.spider.downloader.download(head_req, info.spider) d.addCallback(handle_download_response) d.addErrback(handle_failure) return d class ImageDownload(FilesPipeline): def __init__(self, store_uri, download_func=None, settings=None): self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended super().__init__(store_uri, download_func, settings)
关键说明
stat_file必须返回Deferred对象,适配Twisted异步模型,不能用async def- 两种方案均为非阻塞调用,不会阻塞Scrapy的事件循环
- 方案2更贴合Scrapy生态,能复用现有下载配置,避免重复设置HTTP参数
内容的提问来源于stack exchange,提问作者frenzy
相关产品推荐
相关产品推荐

