You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多页爬取无输出问题求助:代码错误排查

解决Scrapy多页爬取UNESCO表格无输出的问题

我来帮你排查下多页爬取无输出的问题~先对比你的两段代码,单页爬取正常,但换成CrawlSpider后没输出,主要是两个核心原因:

问题分析

1. 第一页数据完全没被处理

CrawlSpider框架会把start_urls的请求默认交给内置的parse方法处理,但你没有重写这个方法来解析第一页的表格,也没让它触发后续的链接提取逻辑——相当于第一页的请求发出去后,框架只是走了默认流程,根本没调用你的parse_table方法。

2. 下一页链接无法被LinkExtractor捕获

UN数据网站的「下一页」按钮基本都不是普通的<a>标签,而是通过表单POST提交来加载新页面的(毕竟是ASP.NET架构的老站点)。而LinkExtractor默认只会提取<a>标签的href属性,你的Rule相当于在找不存在的链接,自然不会发起后续爬取请求。

修正方案

更稳妥的方式是放弃CrawlSpider,回到普通Spider手动处理翻页逻辑,这样灵活性更高。下面分两种情况给出修正代码:

情况1:下一页是普通链接(少数情况)

如果下一页是带href的<a>标签,用这个代码即可,先解析第一页,再递归爬取下一页:

import scrapy
from scrapy.http.request import Request
from indicators.items import EducationIndicators

class mySpider(scrapy.Spider):
    name = "education3"
    allowed_domains = ["data.un.org"]
    start_urls = (
        'http://data.un.org/Data.aspx?d=UNESCO&f=series%3ANER_1',
    )

    def parse(self, response):
        # 先解析当前页面的表格数据
        yield from self.parse_table(response)
        
        # 提取下一页链接(根据实际xpath调整)
        next_page_href = response.xpath('//*[@id="linkNextB"]/@href').extract_first()
        if next_page_href:
            # 处理相对链接,生成完整URL
            next_full_url = response.urljoin(next_page_href)
            yield Request(next_full_url, callback=self.parse)

    def parse_table(self, response):
        sel = response.selector
        for tr in sel.xpath('//*[@id="divData"]/div/table/tr'):
            item = EducationIndicators()
            item['country'] = tr.xpath('td[1]/text()').extract_first()
            item['years'] = tr.xpath('td[position()>1]/text()').extract()
            print(item)
            yield item

情况2:下一页是表单提交(UN站点常见情况)

如果下一页是点击按钮提交表单加载的,需要提取ASP.NET的表单参数(__VIEWSTATE这类),构造POST请求:

import scrapy
from scrapy.http.request import FormRequest
from indicators.items import EducationIndicators

class mySpider(scrapy.Spider):
    name = "education3"
    allowed_domains = ["data.un.org"]
    start_urls = (
        'http://data.un.org/Data.aspx?d=UNESCO&f=series%3ANER_1',
    )

    def parse(self, response):
        # 先解析当前页面的表格
        yield from self.parse_table(response)
        
        # 提取ASP.NET表单的必要参数
        viewstate = response.xpath('//input[@id="__VIEWSTATE"]/@value').extract_first()
        viewstate_gen = response.xpath('//input[@id="__VIEWSTATEGENERATOR"]/@value').extract_first()
        event_validation = response.xpath('//input[@id="__EVENTVALIDATION"]/@value').extract_first()
        
        # 检查下一页按钮是否可用(没有disabled属性)
        next_button = response.xpath('//*[@id="linkNextB"]')
        if next_button and not next_button.xpath('@disabled'):
            # 构造POST请求,模拟点击下一页
            yield FormRequest(
                url=response.url,
                formdata={
                    '__VIEWSTATE': viewstate,
                    '__VIEWSTATEGENERATOR': viewstate_gen,
                    '__EVENTVALIDATION': event_validation,
                    'linkNextB': 'Next >'  # 按钮的value值,需和页面HTML保持一致
                },
                callback=self.parse
            )

    def parse_table(self, response):
        sel = response.selector
        for tr in sel.xpath('//*[@id="divData"]/div/table/tr'):
            item = EducationIndicators()
            item['country'] = tr.xpath('td[1]/text()').extract_first()
            item['years'] = tr.xpath('td[position()>1]/text()').extract()
            print(item)
            yield item

额外提示

  • 你可以打开浏览器的开发者工具(F12),查看下一页按钮的点击请求,确认是GET还是POST,以及需要携带的参数。
  • 如果坚持用CrawlSpider,必须重写parse方法来处理第一页,或者在Rule中添加allow参数匹配start_urls的地址,让第一页也被parse_table处理,但这种方式对表单提交的翻页依然无效。

内容的提问来源于stack exchange,提问作者Roberta Gimenez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:45:36