You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多页爬取数据被覆盖仅返回最后一页问题解决方法

问题原因

代码存在3处核心错误,直接导致仅返回最后一页数据:

  • 在类定义层级直接写分页循环,且在循环内重复定义start_requests、parse_item、parse_book类方法。Python类定义过程中,后定义的同名方法会直接覆盖前序方法,循环结束后仅保留最后一组分页参数对应的方法逻辑,只会发起最后一页的请求。
  • 分页初始值计算错误:初始k=1,第一次循环就执行k +=10得到11,直接跳过第一页(第一页start_count参数应为1)。
  • 列表页xpath逻辑错误:提取表单action属性时使用全局匹配//form,没有基于当前遍历的行节点做相对匹配,会固定拿到页面第一个表单的地址,容易生成错误的详情页跳转链接。
修复后完整代码
import scrapy
from scrapy import FormRequest
from scrapy.crawler import CrawlerProcess
from scrapy.http import Request


class TestSpider(scrapy.Spider):
    name = 'test'
    url = 'https://www.benrishi-navi.com/english/english1_2.php'
    # 公共请求头抽为类属性,无需重复定义
    headers = {
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
        'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7',
        'Cache-Control': 'max-age=0',
        'Connection': 'keep-alive',
        'Content-Type': 'application/x-www-form-urlencoded',
        'Cookie': 'CAKEPHP=u6u40lefkqnm45j49a5i0h6bs3; __utma=42336182.871903078.1657200864.1657200864.1657200864.1; __utmz=42336182.1657200864.1.1.utmcsr=(direct)|utmccn=(direct)|utmcmd=(none)',
        'Origin': 'https://www.benrishi-navi.com',
        'Referer': 'https://www.benrishi-navi.com/english/english1_2.php',
        'Sec-Fetch-Dest': 'document',
        'Sec-Fetch-Mode': 'navigate',
        'Sec-Fetch-Site': 'same-origin',
        'Sec-Fetch-User': '?1',
        'Upgrade-Insecure-Requests': '1',
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36',
        'sec-ch-ua': '".Not/A)Brand";v="99", "Google Chrome";v="103", "Chromium";v="103"',
        'sec-ch-ua-mobile': '?0',
        'sec-ch-ua-platform': '"Windows"'
    }

    def start_requests(self):
        # 分页逻辑放到请求方法内生成,避免类方法被覆盖
        for page_idx in range(5):
            start_count = 1 + page_idx * 10
            search_default = 10 + page_idx * 10
            payload = f'tuusan_year=&tuusan_month=&tuusan_chk=&methodAndOr1=&methodAndOr2=&methodAndOr3=&text_sen=&text_skill=&text_business=&tokkyo_data=&fuki_day_chk=&shuju=&kensyuu_bunya=&text_kensyuu=&methodAndOr_kensyuu=&keitai_kikan=&keitai_hisu=&display_flag=1&search=2&text=&method=&methodAndOr=&area=&pref=&name=&kana=&id=&year=&month=&day=&day_chk=&exp01=&exp02=&exp03=&trip=&venture_support=&venture_flag=&university_support=&university_flag=&university1=&university2=&university=&college=&high_pref=&junior_pref=&elementary_pref=&tyosaku=&hp=&jukoureki=&experience1=&experience2=&experience3=&experience4=&sort=&fuki_year=&fuki_month=&fuki_day=&fuki_day_chk=&id_chk=&shugyou=&fuki=&address1=&address2=&trip_pref=&expref=&office=&max_count=1438&search_count=10&start_count={start_count}&search_default={search_default}'
            yield scrapy.FormRequest(
                url=self.url,
                method='POST',
                body=payload,
                headers=self.headers,
                callback=self.parse_item,
            )

    def parse_item(self, response):
        base_url = "https://www.benrishi-navi.com/english/"
        links = response.xpath("//table[4]//tr")
        for link in links[1:]:
            # 修正xpath为相对路径,匹配当前行下的表单
            t = link.xpath(".//form//@action").get()
            u = link.xpath(".//input[@name='serial']//@value").get()
            if not t or not u:
                continue
            product = base_url + t + "?serial=" + u + "&office_serial=&submit2=Details"
            yield Request(product, callback=self.parse_book, headers=self.headers)

    def parse_book(self, response):
        name = response.xpath("normalize-space(//td[text()[contains(.,'Name')]]/following-sibling::td//text())").get()
        telephone = response.xpath("normalize-space(//td[text()[contains(.,'TEL')]]/following-sibling::td//text())").get()
        fax = response.xpath("normalize-space(//td[text()[contains(.,'FAX')]]/following-sibling::td//text())").get()
        email = response.xpath("normalize-space(//td[text()[contains(.,'Email')]]/following-sibling::td//text())").get()
        website = response.xpath("//td[text()[contains(.,'Website')]]/following-sibling::td//a[starts-with(@href, 'http')]/@href").get()
        registration_date = response.xpath("normalize-space(//td[text()[contains(.,'Registration date')]]/following-sibling::td//text())").get()
        firm = response.xpath("normalize-space(//td[text()[contains(.,'Firm Name')]]/following-sibling::td//text())").get()
        address = response.xpath("normalize-space(//td[text()[contains(.,'Address (Prefecture)')]]/following-sibling::td//text())").get()
        spec = response.xpath("normalize-space(//td[text()[contains(.,'Specialization')]]/following-sibling::td//text())").get()
        if spec:
            spec = spec.replace(" |", "|")
        tech = response.xpath("normalize-space(//td[text()[contains(.,'Technical field')]]/following-sibling::td//text())").get()
        if tech:
            tech = tech.replace(" |", "|")

        yield {
            "name": name,
            "Telephone": telephone,
            "Fax": fax,
            "Email": email,
            "website": website,
            "Registration_date": registration_date,
            "Firm_name": firm,
            "Address": address,
            "Specialization": spec,
            "Technical_field": tech
        }

if __name__ == "__main__":
    process = CrawlerProcess(settings={
        "FEED_URI": "result.json",
        "FEED_FORMAT": "json",
        "CONCURRENT_REQUESTS": 2
    })
    process.crawl(TestSpider)
    process.start()
修复点说明
  • 移除类定义层级的分页循环,将分页参数生成逻辑移到start_requests方法内,为每一页生成独立的POST请求,彻底解决方法覆盖问题。
  • 修正分页参数初始值,从第一页start_count=1开始爬取,不会漏爬首页数据。
  • 修正详情链接提取的xpath逻辑,使用相对路径匹配当前行下的表单地址,避免生成错误跳转链接。
  • 增加空值判断,避免xpath未匹配到内容时调用replace方法抛出异常。
  • 增加并发数配置,降低请求频率避免被站点拦截。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 23:06:08