You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy的FSFilesStore的stat_file方法中实现非阻塞HTTP请求?

问题与解决方案

问题背景

  • 需求:通过HEAD请求获取远程文件的大小、最后修改时间,判断是否需要更新文件,避免下载大体积视频文件
  • 遇到的问题:扩展FSFilesStore重写stat_file方法时,使用async/await调用treq出现连接丢失错误,推测该方法不支持async语法
  • 痛点:现有方案依赖爬虫实现,过于hack,希望在stat_file中用非阻塞方式实现HTTP调用,不侵入爬虫逻辑

代码示例(原错误实现)

import treq
from scrapy.pipelines.files import FilesPipeline
from scrapy.pipelines.files import FSFilesStore

class FSFilesStoreDBCacheExtended(FSFilesStore):

    async def stat_file(self, path, info):
        try:
            head_resp = await treq.head('https://google.com')
        except Exception as e:
            print(e)  # ConnectError here
            raise e

        size = int(head_resp.headers['content-length'])
        mod_timestamp = head_resp.headers['last-modified']
        # ... do some stuff

        return super().stat_file(path, info)


class ImageDownload(FilesPipeline):

    def __init__(self, store_uri, download_func=None, settings=None):
        self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended
        self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended
        super().__init__(store_uri, download_func, settings)

错误信息(中文翻译)

[<twisted.python.failure.Failure twisted.internet.error.ConnectionLost: 连接已非正常断开: 连接丢失。>]

解决方案

Scrapy基于Twisted异步框架,stat_file方法需要返回Deferred对象而非使用async/await(直接用async def会导致事件循环不兼容)。以下提供两种可行方案:

方案1:使用treq(基于Twisted)实现非阻塞HEAD请求

利用treq返回Deferred的特性,通过回调链处理响应和异常:

import treq
from scrapy.pipelines.files import FilesPipeline, FSFilesStore

class FSFilesStoreDBCacheExtended(FSFilesStore):

    def stat_file(self, path, info):
        # 从info中获取远程文件URL(FilesPipeline会传递该元数据)
        remote_url = info.get('url')
        if not remote_url:
            return super().stat_file(path, info)

        def handle_head_response(response):
            # 解析HEAD响应头
            size = int(response.headers.get('content-length', 0))
            mod_timestamp = response.headers.get('last-modified')
            
            # 此处添加你的更新判断逻辑:
            # 1. 对比本地文件与远程文件的大小/修改时间
            # 2. 若无需更新,可直接返回自定义状态跳过下载
            
            # 执行父类默认逻辑(或根据判断结果返回对应状态)
            return super(FSFilesStoreDBCacheExtended, self).stat_file(path, info)

        def handle_failure(failure):
            # 请求失败时回退到默认逻辑
            print(f"HEAD请求失败: {failure.getErrorMessage()}")
            return super(FSFilesStoreDBCacheExtended, self).stat_file(path, info)

        # 发送HEAD请求并返回Deferred
        d = treq.head(remote_url)
        d.addCallback(handle_head_response)
        d.addErrback(handle_failure)
        return d

class ImageDownload(FilesPipeline):

    def __init__(self, store_uri, download_func=None, settings=None):
        self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended
        self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended
        super().__init__(store_uri, download_func, settings)

方案2:复用Scrapy下载器发送HEAD请求(推荐)

直接使用Scrapy内置的下载器,可复用爬虫的UA、代理、重试策略等配置:

from scrapy.pipelines.files import FilesPipeline, FSFilesStore
from scrapy.http import Request

class FSFilesStoreDBCacheExtended(FSFilesStore):

    def stat_file(self, path, info):
        remote_url = info.get('url')
        if not remote_url:
            return super().stat_file(path, info)

        def handle_download_response(response):
            size = int(response.headers.get('content-length', 0))
            mod_timestamp = response.headers.get('last-modified')
            
            # 你的更新判断逻辑
            
            return super().stat_file(path, info)

        def handle_failure(failure):
            print(f"HEAD请求失败: {failure.getErrorMessage()}")
            return super().stat_file(path, info)

        # 构造HEAD请求,通过Scrapy下载器发送
        head_req = Request(remote_url, method='HEAD')
        d = info.spider.downloader.download(head_req, info.spider)
        d.addCallback(handle_download_response)
        d.addErrback(handle_failure)
        return d

class ImageDownload(FilesPipeline):

    def __init__(self, store_uri, download_func=None, settings=None):
        self.STORE_SCHEMES[''] = FSFilesStoreDBCacheExtended
        self.STORE_SCHEMES['file'] = FSFilesStoreDBCacheExtended
        super().__init__(store_uri, download_func, settings)

关键说明

  • stat_file必须返回Deferred对象,适配Twisted异步模型,不能用async def
  • 两种方案均为非阻塞调用,不会阻塞Scrapy的事件循环
  • 方案2更贴合Scrapy生态,能复用现有下载配置,避免重复设置HTTP参数

内容的提问来源于stack exchange,提问作者frenzy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 10:45:04