You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取重定向PDF时触发AttributeError错误求助

解决方案

1. 在Parse方法中直接过滤非文本响应

在parse方法最开始做校验,跳过PDF这类二进制响应的处理:

def parse(self, response):
    # 先判断是否为文本响应,非文本直接跳过
    try:
        response.text
    except AttributeError:
        return

    internal_le = LinkExtractor(
        allow_domains=tld_t,
        unique=True,
    )
    in_links = internal_le.extract_links(response)

    for link in in_links:
        if link.url:
            yield Request(
                link.url,
                callback=self.parse,
            )

2. 自定义下载中间件拦截PDF响应

从下载阶段直接拦截PDF内容,彻底避免后续报错:

在项目的middlewares.py中添加以下中间件:

from scrapy.exceptions import IgnoreRequest

class BlockPDFMiddleware:
    def process_response(self, request, response, spider):
        # 双重校验:先检查Content-Type,再验证URL后缀
        content_type = response.headers.get('Content-Type', b'').decode('utf-8', errors='ignore')
        if 'application/pdf' in content_type or response.url.lower().endswith('.pdf'):
            raise IgnoreRequest(f"跳过PDF内容: {response.url}")
        return response

然后在settings.py里启用该中间件(优先级设为543,确保在默认重定向中间件之后执行):

DOWNLOADER_MIDDLEWARES = {
    '你的项目名.middlewares.BlockPDFMiddleware': 543,
}

3. 提前过滤可能重定向到PDF的链接(可选)

如果能预判部分内部链接会跳转到PDF,可以给LinkExtractor添加规则过滤这类路径:

internal_le = LinkExtractor(
    allow_domains=tld_t,
    unique=True,
    deny=r'(download|document|pdf)\.php$'  # 根据网站实际路径调整正则规则
)

内容的提问来源于stack exchange,提问作者leeprevost

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 01:17:28