Scrapy无法正常跳转URL,爬取quotes.toscrape.com仅能获取第一页如何解决?
Scrapy仅爬取第一页问题修复方案
核心问题及修复点
allowed_domains配置不符合规范allowed_domains仅接收纯域名,不需要携带路径后缀,你配置中末尾的/会触发Scrapy的跨域拦截规则,所有分页请求会被直接过滤,不会发出。
修复后配置:allowed_domains = ['quotes.toscrape.com']
- 分页请求代码缩进错误
你把下一页请求的逻辑写在了遍历单条quote的for循环内部,会导致两个问题:一是每处理一条quote就重复发起一次下一页请求,产生大量无效请求;二是如果页面quote为空,永远不会触发分页请求。
修复方法:将下一页相关代码整体移出for循环,取消一层缩进。
- 分页请求代码缩进错误
- 缺少下一页链接判空逻辑
最后一页不存在下一页按钮时,next_page_url会返回None,直接调用urljoin会抛出异常,需要增加非空判断再发起请求。
- 缺少下一页链接判空逻辑
修复后完整代码
import scrapy class QuotesSpider(scrapy.Spider): name = 'quotes' allowed_domains = ['quotes.toscrape.com'] start_urls = ['http://quotes.toscrape.com/'] def parse(self, response): quotes = response.xpath('//*[@class="quote"]') for quote in quotes : text = quote.xpath('.//*[@class="text"]/text()').extract_first() author = quote.xpath('.//*[@class="author"]/text()').extract_first() tags = quote.xpath('.//*[@class="keywords"]/@content').extract_first() yield{ 'text':text, 'author':author, 'tags':tags } next_page_url = response.xpath('//*[@class="next"]/a/@href').extract_first() if next_page_url: absolute_next_page_url = response.urljoin(next_page_url) yield scrapy.Request(absolute_next_page_url)
修改完成后重新运行即可正常爬取全站10页的全部名言数据。
内容的提问来源于stack exchange,提问作者Faycal Faycal
相关产品推荐
相关产品推荐

