You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫twisted.python.failure.Failure错误的永久解决方案

问题根因

你遇到的twisted.python.failure.Failure twisted.internet连接错误并非单纯由请求频率过高导致,仅设置download_delay无法彻底解决,核心诱因有4点:

  • 硬编码的Cookie已过期:代码中写死的CAKEPHP、__utma等字段是单次访问生成的临时会话凭证,过期后站点会直接断开TCP连接,与请求间隔无关
  • Xpath逻辑存在bug:parse_item方法中提取表单action的路径使用了全局//form匹配,未绑定当前循环的tr节点,会生成大量无效重复的详情页请求,触发站点反爬拦截
  • 请求头配置不兼容:手动写死Connection: keep-alive,但Scrapy默认HTTP下载器不支持长连接复用,会触发协议层连接重置
  • 缺乏容错机制:默认配置下遇到连接错误直接抛出失败,无自动退避重试逻辑,偶发网络波动也会直接中断任务
永久修复方案

按以下规则调整即可保障任务稳定运行:

  • 删除硬编码的Cookie配置,启用Scrapy自带Cookie中间件自动维护会话,无需手动维护临时凭证
  • 修正Xpath为相对路径匹配,提取当前行下的表单参数,避免生成无效请求
  • 移除请求头中硬编码的Connection字段,由下载器自动适配连接规则
  • 增加并发控制、随机下载延迟、错误重试配置,模拟真人访问行为,规避反爬拦截
  • 动态提取列表页总条数构造表单参数,替代硬编码的固定计数值,避免参数不匹配被站点拒绝
修正后可稳定运行的代码
import scrapy
from scrapy import FormRequest
from scrapy.crawler import CrawlerProcess
from scrapy.http import Request


class TestSpider(scrapy.Spider):
    name = 'test'
    start_url = 'https://www.benrishi-navi.com/english/english1_2.php'
    custom_settings = {
        'CONCURRENT_REQUESTS': 1,
        'DOWNLOAD_DELAY': 3,
        'RANDOMIZE_DOWNLOAD_DELAY': True,
        'RETRY_TIMES': 5,
        'RETRY_HTTP_CODES': [500, 502, 503, 504, 403, 408, 429],
        'RETRY_PRIORITY_ADJUST': -1,
        'COOKIES_ENABLED': True,
        'REDIRECT_ENABLED': True,
        'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
    }

    headers = {
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
        'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7',
        'Cache-Control': 'max-age=0',
        'Origin': 'https://www.benrishi-navi.com',
        'Referer': 'https://www.benrishi-navi.com/english/english1_2.php',
        'Sec-Fetch-Dest': 'document',
        'Sec-Fetch-Mode': 'navigate',
        'Sec-Fetch-Site': 'same-origin',
        'Sec-Fetch-User': '?1',
        'Upgrade-Insecure-Requests': '1',
        'sec-ch-ua': '"Not_A Brand";v="8", "Chromium";v="120", "Google Chrome";v="120"',
        'sec-ch-ua-mobile': '?0',
        'sec-ch-ua-platform': '"Windows"',
    }

    def get_base_formdata(self):
        return [
            ('tuusan_year', ''),
            ('tuusan_month', ''),
            ('tuusan_chk', ''),
            ('methodAndOr1', ''),
            ('methodAndOr2', ''),
            ('methodAndOr3', ''),
            ('text_sen', ''),
            ('text_skill', ''),
            ('text_business', ''),
            ('tokkyo_data', ''),
            ('fuki_day_chk', ''),
            ('shuju', ''),
            ('kensyuu_bunya', ''),
            ('text_kensyuu', ''),
            ('methodAndOr_kensyuu', ''),
            ('keitai_kikan', ''),
            ('keitai_hisu', ''),
            ('display_flag', '1'),
            ('search', '2'),
            ('text', ''),
            ('method', ''),
            ('methodAndOr', ''),
            ('area', ''),
            ('pref', ''),
            ('name', ''),
            ('kana', ''),
            ('id', ''),
            ('year', ''),
            ('month', ''),
            ('day', ''),
            ('day_chk', ''),
            ('exp01', ''),
            ('exp02', ''),
            ('exp03', ''),
            ('trip', ''),
            ('venture_support', ''),
            ('venture_flag', ''),
            ('university_support', ''),
            ('university_flag', ''),
            ('university1', ''),
            ('university2', ''),
            ('university', ''),
            ('college', ''),
            ('high_pref', ''),
            ('junior_pref', ''),
            ('elementary_pref', ''),
            ('tyosaku', ''),
            ('hp', ''),
            ('jukoureki', ''),
            ('experience1', ''),
            ('experience2', ''),
            ('experience3', ''),
            ('experience4', ''),
            ('sort', ''),
            ('fuki_year', ''),
            ('fuki_month', ''),
            ('fuki_day', ''),
            ('fuki_day_chk', ''),
            ('id_chk', ''),
            ('shugyou', ''),
            ('fuki', ''),
            ('address1', ''),
            ('address2', ''),
            ('trip_pref', ''),
            ('expref', ''),
            ('office', ''),
            ('start_count', '1'),
            ('search_default', '1000'),
        ]
   
    
    def start_requests(self):
        yield Request(
            url=self.start_url,
            headers=self.headers,
            callback=self.submit_search
        )
    
    def submit_search(self, response):
        total_count = response.xpath("//input[@name='max_count']/@value").get()
        formdata = self.get_base_formdata()
        formdata.append(('max_count', total_count if total_count else '1438'))
        formdata.append(('search_count', total_count if total_count else '1438'))
        
        yield FormRequest.from_response(
            response,
            method='POST',
            formdata=dict(formdata),
            headers=self.headers,
            callback=self.parse_item,
            dont_filter=True
        )
        
        
    def parse_item(self, response):
        base_url = "https://www.benrishi-navi.com/english/"
        links = response.xpath("//table[4]//tr")
        form_action = response.xpath("//form[1]/@action").get()
        for link in links[1:]:
            u = link.xpath(".//input[@name='serial']/@value").get()
            if not u or not form_action:
                continue
            product = f"{base_url}{form_action}?serial={u}&office_serial=&submit2=Details"
            yield Request(product, callback=self.parse_book, headers=self.headers)
                    
    def parse_book(self,response):
        name=response.xpath("normalize-space(//td[text()[contains(.,'Name')]]/following-sibling::td//text())").get()
        telephone=response.xpath("normalize-space(//td[text()[contains(.,'TEL')]]/following-sibling::td//text())").get()
        fax=response.xpath("normalize-space(//td[text()[contains(.,'FAX')]]/following-sibling::td//text())").get()
        email=response.xpath("normalize-space(//td[text()[contains(.,'Email')]]/following-sibling::td//text())").get()
        website=response.xpath("//td[text()[contains(.,'Website')]]/following-sibling::td//a[starts-with(@href, 'http')]/@href").get()
        registration_date=response.xpath("normalize-space(//td[text()[contains(.,'Registration date')]]/following-sibling::td//text())").get()
        firm=response.xpath("normalize-space(//td[text()[contains(.,'Firm Name')]]/following-sibling::td//text())").get()
        address=response.xpath("normalize-space(//td[text()[contains(.,'Address (Prefecture)')]]/following-sibling::td//text())").get()
        spec_raw = response.xpath("normalize-space(//td[text()[contains(.,'Specialization')]]/following-sibling::td//text())").get()
        tech_raw = response.xpath("normalize-space(//td[text()[contains(.,'Technical field')]]/following-sibling::td//text())").get()
        spec = spec_raw.replace(" |","|") if spec_raw else ""
        tech = tech_raw.replace(" |","|") if tech_raw else ""
        
        yield{
            "name":name,
            "Telephone":telephone,
            "Fax":fax,
            "Email":email,
            "website":website,
            "Registration_date":registration_date,
            "Firm_name":firm,
            "Address":address,
            "Specialization":spec,
            "Technical_field":tech
        }

if __name__ == "__main__":
    process = CrawlerProcess(settings={
        'FEED_URI': 'patent_agents.csv',
        'FEED_FORMAT': 'csv'
    })
    process.crawl(TestSpider)
    process.start()
配置说明
  • 开启RANDOMIZE_DOWNLOAD_DELAY后,实际请求间隔会在设置的DOWNLOAD_DELAY值的0.5-1.5倍之间随机波动,比固定延迟更难被反爬识别
  • 重试配置会在触发连接错误、访问受限状态码时自动退避重试,不会因为单次网络波动中断整个任务
  • 先访问首页获取有效Cookie再提交搜索表单,完全模拟真人访问路径,不会出现凭证过期被拦截的问题
  • 增加空值判断,避免因为页面结构微小变动导致代码抛出异常中断

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:06:24