You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy自定义文件管道获None类型response,无法下载文件

问题:Scrapy自定义文件管道时response为None导致文件下载失败

启用默认FilesPipeline时,HTML/PDF文件均可正常下载,但自定义继承FilesPipeline的MyCustomFilePipeline并重写file_path方法后,出现以下问题:

  • 控制台打印前两条调试信息,但response为<class 'NoneType'>
  • 后续打印语句未执行,文件未下载,files字段未填充

自定义管道代码如下:

class MyCustomFilePipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, *, item=None):        
        # extract from 191148 http://mywebsite.com/filedownload.asp?pn=191148&yr=2022
        pn = re.search(r'(?<=pn\=)\d+', request.url).group()
        print(f'{request.url} - {pn}')
        print(type(response)) # <-- this prints as <class 'NoneType'>
        
        response_contentype = response.headers['Content-Type'].decode('ASCII')
        ext = 'html'
        if response_contentype  == 'text/html':
            ext = 'html'
        elif response_contentype == 'application/pdf':
            ext = 'pdf'
        print(f'{pn}.{ext}') # <-- this is not printed  
        return f'{pn}.{ext}'

解决方法

问题根源

Scrapy的FilesPipeline中,file_path方法会被调用两次:

  1. 第一次是在准备发起下载请求前,此时还没有响应,所以response为None,目的是预先生成文件名占位
  2. 第二次是在下载完成后,此时会传入完整的response,用于最终确定文件名

你的代码直接访问response.headers,在第一次调用时会触发AttributeError,导致管道中断,后续下载流程无法执行,所以文件没下载、files字段也没填充。

方案1:兼容response为None的情况

在代码里先判断response是否存在,不存在时返回临时文件名,等下载完成后再返回最终文件名:

import re
from scrapy.pipelines.files import FilesPipeline

class MyCustomFilePipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, *, item=None):        
        pn = re.search(r'(?<=pn\=)\d+', request.url).group()
        print(f'{request.url} - {pn}')
        
        # 处理第一次调用时response为None的情况
        if not response:
            # 返回临时后缀,后续会被覆盖
            return f'{pn}.tmp'
            
        response_contentype = response.headers['Content-Type'].decode('ASCII')
        ext = 'html'
        if response_contentype == 'application/pdf':
            ext = 'pdf'
        print(f'{pn}.{ext}')
        return f'{pn}.{ext}'

方案2:提前传递文件类型(更可靠)

在爬虫阶段就解析出文件类型,放到请求的meta中,这样file_path方法不用依赖response就能确定后缀:

爬虫代码示例

import scrapy

class MySpider(scrapy.Spider):
    name = 'my_spider'
    start_urls = ['http://mywebsite.com/file_list']
    
    def parse(self, response):
        # 假设从页面解析出文件ID和对应的类型
        file_items = [
            {'pn': '191148', 'type': 'pdf'},
            {'pn': '191149', 'type': 'html'}
        ]
        
        for item in file_items:
            file_url = f'http://mywebsite.com/filedownload.asp?pn={item["pn"]}&yr=2022'
            yield {
                'file_urls': [
                    scrapy.Request(
                        file_url,
                        meta={'file_ext': item['type']}  # 传递文件类型
                    )
                ]
            }

自定义管道代码

import re
from scrapy.pipelines.files import FilesPipeline

class MyCustomFilePipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, *, item=None):        
        pn = re.search(r'(?<=pn\=)\d+', request.url).group()
        # 从request.meta中获取提前传递的文件后缀
        ext = request.meta.get('file_ext', 'html')
        print(f'{pn}.{ext}')
        return f'{pn}.{ext}'

内容的提问来源于stack exchange,提问作者claws

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 21:45:28