Scrapy爬虫twisted.python.failure.Failure错误的永久解决方案
问题根因
你遇到的twisted.python.failure.Failure twisted.internet连接错误并非单纯由请求频率过高导致,仅设置download_delay无法彻底解决,核心诱因有4点:
- 硬编码的Cookie已过期:代码中写死的
CAKEPHP、__utma等字段是单次访问生成的临时会话凭证,过期后站点会直接断开TCP连接,与请求间隔无关 - Xpath逻辑存在bug:
parse_item方法中提取表单action的路径使用了全局//form匹配,未绑定当前循环的tr节点,会生成大量无效重复的详情页请求,触发站点反爬拦截 - 请求头配置不兼容:手动写死
Connection: keep-alive,但Scrapy默认HTTP下载器不支持长连接复用,会触发协议层连接重置 - 缺乏容错机制:默认配置下遇到连接错误直接抛出失败,无自动退避重试逻辑,偶发网络波动也会直接中断任务
永久修复方案
按以下规则调整即可保障任务稳定运行:
- 删除硬编码的Cookie配置,启用Scrapy自带Cookie中间件自动维护会话,无需手动维护临时凭证
- 修正Xpath为相对路径匹配,提取当前行下的表单参数,避免生成无效请求
- 移除请求头中硬编码的
Connection字段,由下载器自动适配连接规则 - 增加并发控制、随机下载延迟、错误重试配置,模拟真人访问行为,规避反爬拦截
- 动态提取列表页总条数构造表单参数,替代硬编码的固定计数值,避免参数不匹配被站点拒绝
修正后可稳定运行的代码
import scrapy from scrapy import FormRequest from scrapy.crawler import CrawlerProcess from scrapy.http import Request class TestSpider(scrapy.Spider): name = 'test' start_url = 'https://www.benrishi-navi.com/english/english1_2.php' custom_settings = { 'CONCURRENT_REQUESTS': 1, 'DOWNLOAD_DELAY': 3, 'RANDOMIZE_DOWNLOAD_DELAY': True, 'RETRY_TIMES': 5, 'RETRY_HTTP_CODES': [500, 502, 503, 504, 403, 408, 429], 'RETRY_PRIORITY_ADJUST': -1, 'COOKIES_ENABLED': True, 'REDIRECT_ENABLED': True, 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } headers = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7', 'Cache-Control': 'max-age=0', 'Origin': 'https://www.benrishi-navi.com', 'Referer': 'https://www.benrishi-navi.com/english/english1_2.php', 'Sec-Fetch-Dest': 'document', 'Sec-Fetch-Mode': 'navigate', 'Sec-Fetch-Site': 'same-origin', 'Sec-Fetch-User': '?1', 'Upgrade-Insecure-Requests': '1', 'sec-ch-ua': '"Not_A Brand";v="8", "Chromium";v="120", "Google Chrome";v="120"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"', } def get_base_formdata(self): return [ ('tuusan_year', ''), ('tuusan_month', ''), ('tuusan_chk', ''), ('methodAndOr1', ''), ('methodAndOr2', ''), ('methodAndOr3', ''), ('text_sen', ''), ('text_skill', ''), ('text_business', ''), ('tokkyo_data', ''), ('fuki_day_chk', ''), ('shuju', ''), ('kensyuu_bunya', ''), ('text_kensyuu', ''), ('methodAndOr_kensyuu', ''), ('keitai_kikan', ''), ('keitai_hisu', ''), ('display_flag', '1'), ('search', '2'), ('text', ''), ('method', ''), ('methodAndOr', ''), ('area', ''), ('pref', ''), ('name', ''), ('kana', ''), ('id', ''), ('year', ''), ('month', ''), ('day', ''), ('day_chk', ''), ('exp01', ''), ('exp02', ''), ('exp03', ''), ('trip', ''), ('venture_support', ''), ('venture_flag', ''), ('university_support', ''), ('university_flag', ''), ('university1', ''), ('university2', ''), ('university', ''), ('college', ''), ('high_pref', ''), ('junior_pref', ''), ('elementary_pref', ''), ('tyosaku', ''), ('hp', ''), ('jukoureki', ''), ('experience1', ''), ('experience2', ''), ('experience3', ''), ('experience4', ''), ('sort', ''), ('fuki_year', ''), ('fuki_month', ''), ('fuki_day', ''), ('fuki_day_chk', ''), ('id_chk', ''), ('shugyou', ''), ('fuki', ''), ('address1', ''), ('address2', ''), ('trip_pref', ''), ('expref', ''), ('office', ''), ('start_count', '1'), ('search_default', '1000'), ] def start_requests(self): yield Request( url=self.start_url, headers=self.headers, callback=self.submit_search ) def submit_search(self, response): total_count = response.xpath("//input[@name='max_count']/@value").get() formdata = self.get_base_formdata() formdata.append(('max_count', total_count if total_count else '1438')) formdata.append(('search_count', total_count if total_count else '1438')) yield FormRequest.from_response( response, method='POST', formdata=dict(formdata), headers=self.headers, callback=self.parse_item, dont_filter=True ) def parse_item(self, response): base_url = "https://www.benrishi-navi.com/english/" links = response.xpath("//table[4]//tr") form_action = response.xpath("//form[1]/@action").get() for link in links[1:]: u = link.xpath(".//input[@name='serial']/@value").get() if not u or not form_action: continue product = f"{base_url}{form_action}?serial={u}&office_serial=&submit2=Details" yield Request(product, callback=self.parse_book, headers=self.headers) def parse_book(self,response): name=response.xpath("normalize-space(//td[text()[contains(.,'Name')]]/following-sibling::td//text())").get() telephone=response.xpath("normalize-space(//td[text()[contains(.,'TEL')]]/following-sibling::td//text())").get() fax=response.xpath("normalize-space(//td[text()[contains(.,'FAX')]]/following-sibling::td//text())").get() email=response.xpath("normalize-space(//td[text()[contains(.,'Email')]]/following-sibling::td//text())").get() website=response.xpath("//td[text()[contains(.,'Website')]]/following-sibling::td//a[starts-with(@href, 'http')]/@href").get() registration_date=response.xpath("normalize-space(//td[text()[contains(.,'Registration date')]]/following-sibling::td//text())").get() firm=response.xpath("normalize-space(//td[text()[contains(.,'Firm Name')]]/following-sibling::td//text())").get() address=response.xpath("normalize-space(//td[text()[contains(.,'Address (Prefecture)')]]/following-sibling::td//text())").get() spec_raw = response.xpath("normalize-space(//td[text()[contains(.,'Specialization')]]/following-sibling::td//text())").get() tech_raw = response.xpath("normalize-space(//td[text()[contains(.,'Technical field')]]/following-sibling::td//text())").get() spec = spec_raw.replace(" |","|") if spec_raw else "" tech = tech_raw.replace(" |","|") if tech_raw else "" yield{ "name":name, "Telephone":telephone, "Fax":fax, "Email":email, "website":website, "Registration_date":registration_date, "Firm_name":firm, "Address":address, "Specialization":spec, "Technical_field":tech } if __name__ == "__main__": process = CrawlerProcess(settings={ 'FEED_URI': 'patent_agents.csv', 'FEED_FORMAT': 'csv' }) process.crawl(TestSpider) process.start()
配置说明
- 开启
RANDOMIZE_DOWNLOAD_DELAY后,实际请求间隔会在设置的DOWNLOAD_DELAY值的0.5-1.5倍之间随机波动,比固定延迟更难被反爬识别 - 重试配置会在触发连接错误、访问受限状态码时自动退避重试,不会因为单次网络波动中断整个任务
- 先访问首页获取有效Cookie再提交搜索表单,完全模拟真人访问路径,不会出现凭证过期被拦截的问题
- 增加空值判断,避免因为页面结构微小变动导致代码抛出异常中断
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

