Scrapy多页爬取数据被覆盖仅返回最后一页问题解决方法
问题原因
代码存在3处核心错误,直接导致仅返回最后一页数据:
- 在类定义层级直接写分页循环,且在循环内重复定义
start_requests、parse_item、parse_book类方法。Python类定义过程中,后定义的同名方法会直接覆盖前序方法,循环结束后仅保留最后一组分页参数对应的方法逻辑,只会发起最后一页的请求。 - 分页初始值计算错误:初始k=1,第一次循环就执行
k +=10得到11,直接跳过第一页(第一页start_count参数应为1)。 - 列表页xpath逻辑错误:提取表单action属性时使用全局匹配
//form,没有基于当前遍历的行节点做相对匹配,会固定拿到页面第一个表单的地址,容易生成错误的详情页跳转链接。
修复后完整代码
import scrapy from scrapy import FormRequest from scrapy.crawler import CrawlerProcess from scrapy.http import Request class TestSpider(scrapy.Spider): name = 'test' url = 'https://www.benrishi-navi.com/english/english1_2.php' # 公共请求头抽为类属性,无需重复定义 headers = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7', 'Cache-Control': 'max-age=0', 'Connection': 'keep-alive', 'Content-Type': 'application/x-www-form-urlencoded', 'Cookie': 'CAKEPHP=u6u40lefkqnm45j49a5i0h6bs3; __utma=42336182.871903078.1657200864.1657200864.1657200864.1; __utmz=42336182.1657200864.1.1.utmcsr=(direct)|utmccn=(direct)|utmcmd=(none)', 'Origin': 'https://www.benrishi-navi.com', 'Referer': 'https://www.benrishi-navi.com/english/english1_2.php', 'Sec-Fetch-Dest': 'document', 'Sec-Fetch-Mode': 'navigate', 'Sec-Fetch-Site': 'same-origin', 'Sec-Fetch-User': '?1', 'Upgrade-Insecure-Requests': '1', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36', 'sec-ch-ua': '".Not/A)Brand";v="99", "Google Chrome";v="103", "Chromium";v="103"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"' } def start_requests(self): # 分页逻辑放到请求方法内生成,避免类方法被覆盖 for page_idx in range(5): start_count = 1 + page_idx * 10 search_default = 10 + page_idx * 10 payload = f'tuusan_year=&tuusan_month=&tuusan_chk=&methodAndOr1=&methodAndOr2=&methodAndOr3=&text_sen=&text_skill=&text_business=&tokkyo_data=&fuki_day_chk=&shuju=&kensyuu_bunya=&text_kensyuu=&methodAndOr_kensyuu=&keitai_kikan=&keitai_hisu=&display_flag=1&search=2&text=&method=&methodAndOr=&area=&pref=&name=&kana=&id=&year=&month=&day=&day_chk=&exp01=&exp02=&exp03=&trip=&venture_support=&venture_flag=&university_support=&university_flag=&university1=&university2=&university=&college=&high_pref=&junior_pref=&elementary_pref=&tyosaku=&hp=&jukoureki=&experience1=&experience2=&experience3=&experience4=&sort=&fuki_year=&fuki_month=&fuki_day=&fuki_day_chk=&id_chk=&shugyou=&fuki=&address1=&address2=&trip_pref=&expref=&office=&max_count=1438&search_count=10&start_count={start_count}&search_default={search_default}' yield scrapy.FormRequest( url=self.url, method='POST', body=payload, headers=self.headers, callback=self.parse_item, ) def parse_item(self, response): base_url = "https://www.benrishi-navi.com/english/" links = response.xpath("//table[4]//tr") for link in links[1:]: # 修正xpath为相对路径,匹配当前行下的表单 t = link.xpath(".//form//@action").get() u = link.xpath(".//input[@name='serial']//@value").get() if not t or not u: continue product = base_url + t + "?serial=" + u + "&office_serial=&submit2=Details" yield Request(product, callback=self.parse_book, headers=self.headers) def parse_book(self, response): name = response.xpath("normalize-space(//td[text()[contains(.,'Name')]]/following-sibling::td//text())").get() telephone = response.xpath("normalize-space(//td[text()[contains(.,'TEL')]]/following-sibling::td//text())").get() fax = response.xpath("normalize-space(//td[text()[contains(.,'FAX')]]/following-sibling::td//text())").get() email = response.xpath("normalize-space(//td[text()[contains(.,'Email')]]/following-sibling::td//text())").get() website = response.xpath("//td[text()[contains(.,'Website')]]/following-sibling::td//a[starts-with(@href, 'http')]/@href").get() registration_date = response.xpath("normalize-space(//td[text()[contains(.,'Registration date')]]/following-sibling::td//text())").get() firm = response.xpath("normalize-space(//td[text()[contains(.,'Firm Name')]]/following-sibling::td//text())").get() address = response.xpath("normalize-space(//td[text()[contains(.,'Address (Prefecture)')]]/following-sibling::td//text())").get() spec = response.xpath("normalize-space(//td[text()[contains(.,'Specialization')]]/following-sibling::td//text())").get() if spec: spec = spec.replace(" |", "|") tech = response.xpath("normalize-space(//td[text()[contains(.,'Technical field')]]/following-sibling::td//text())").get() if tech: tech = tech.replace(" |", "|") yield { "name": name, "Telephone": telephone, "Fax": fax, "Email": email, "website": website, "Registration_date": registration_date, "Firm_name": firm, "Address": address, "Specialization": spec, "Technical_field": tech } if __name__ == "__main__": process = CrawlerProcess(settings={ "FEED_URI": "result.json", "FEED_FORMAT": "json", "CONCURRENT_REQUESTS": 2 }) process.crawl(TestSpider) process.start()
修复点说明
- 移除类定义层级的分页循环,将分页参数生成逻辑移到
start_requests方法内,为每一页生成独立的POST请求,彻底解决方法覆盖问题。 - 修正分页参数初始值,从第一页
start_count=1开始爬取,不会漏爬首页数据。 - 修正详情链接提取的xpath逻辑,使用相对路径匹配当前行下的表单地址,避免生成错误跳转链接。
- 增加空值判断,避免xpath未匹配到内容时调用
replace方法抛出异常。 - 增加并发数配置,降低请求频率避免被站点拦截。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

