Scrapy多页爬取无输出问题求助:代码错误排查
解决Scrapy多页爬取UNESCO表格无输出的问题
我来帮你排查下多页爬取无输出的问题~先对比你的两段代码,单页爬取正常,但换成CrawlSpider后没输出,主要是两个核心原因:
问题分析
1. 第一页数据完全没被处理
CrawlSpider框架会把start_urls的请求默认交给内置的parse方法处理,但你没有重写这个方法来解析第一页的表格,也没让它触发后续的链接提取逻辑——相当于第一页的请求发出去后,框架只是走了默认流程,根本没调用你的parse_table方法。
2. 下一页链接无法被LinkExtractor捕获
UN数据网站的「下一页」按钮基本都不是普通的<a>标签,而是通过表单POST提交来加载新页面的(毕竟是ASP.NET架构的老站点)。而LinkExtractor默认只会提取<a>标签的href属性,你的Rule相当于在找不存在的链接,自然不会发起后续爬取请求。
修正方案
更稳妥的方式是放弃CrawlSpider,回到普通Spider手动处理翻页逻辑,这样灵活性更高。下面分两种情况给出修正代码:
情况1:下一页是普通链接(少数情况)
如果下一页是带href的<a>标签,用这个代码即可,先解析第一页,再递归爬取下一页:
import scrapy from scrapy.http.request import Request from indicators.items import EducationIndicators class mySpider(scrapy.Spider): name = "education3" allowed_domains = ["data.un.org"] start_urls = ( 'http://data.un.org/Data.aspx?d=UNESCO&f=series%3ANER_1', ) def parse(self, response): # 先解析当前页面的表格数据 yield from self.parse_table(response) # 提取下一页链接(根据实际xpath调整) next_page_href = response.xpath('//*[@id="linkNextB"]/@href').extract_first() if next_page_href: # 处理相对链接,生成完整URL next_full_url = response.urljoin(next_page_href) yield Request(next_full_url, callback=self.parse) def parse_table(self, response): sel = response.selector for tr in sel.xpath('//*[@id="divData"]/div/table/tr'): item = EducationIndicators() item['country'] = tr.xpath('td[1]/text()').extract_first() item['years'] = tr.xpath('td[position()>1]/text()').extract() print(item) yield item
情况2:下一页是表单提交(UN站点常见情况)
如果下一页是点击按钮提交表单加载的,需要提取ASP.NET的表单参数(__VIEWSTATE这类),构造POST请求:
import scrapy from scrapy.http.request import FormRequest from indicators.items import EducationIndicators class mySpider(scrapy.Spider): name = "education3" allowed_domains = ["data.un.org"] start_urls = ( 'http://data.un.org/Data.aspx?d=UNESCO&f=series%3ANER_1', ) def parse(self, response): # 先解析当前页面的表格 yield from self.parse_table(response) # 提取ASP.NET表单的必要参数 viewstate = response.xpath('//input[@id="__VIEWSTATE"]/@value').extract_first() viewstate_gen = response.xpath('//input[@id="__VIEWSTATEGENERATOR"]/@value').extract_first() event_validation = response.xpath('//input[@id="__EVENTVALIDATION"]/@value').extract_first() # 检查下一页按钮是否可用(没有disabled属性) next_button = response.xpath('//*[@id="linkNextB"]') if next_button and not next_button.xpath('@disabled'): # 构造POST请求,模拟点击下一页 yield FormRequest( url=response.url, formdata={ '__VIEWSTATE': viewstate, '__VIEWSTATEGENERATOR': viewstate_gen, '__EVENTVALIDATION': event_validation, 'linkNextB': 'Next >' # 按钮的value值,需和页面HTML保持一致 }, callback=self.parse ) def parse_table(self, response): sel = response.selector for tr in sel.xpath('//*[@id="divData"]/div/table/tr'): item = EducationIndicators() item['country'] = tr.xpath('td[1]/text()').extract_first() item['years'] = tr.xpath('td[position()>1]/text()').extract() print(item) yield item
额外提示
- 你可以打开浏览器的开发者工具(F12),查看下一页按钮的点击请求,确认是GET还是POST,以及需要携带的参数。
- 如果坚持用
CrawlSpider,必须重写parse方法来处理第一页,或者在Rule中添加allow参数匹配start_urls的地址,让第一页也被parse_table处理,但这种方式对表单提交的翻页依然无效。
内容的提问来源于stack exchange,提问作者Roberta Gimenez
相关产品推荐
相关产品推荐

