You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy无法正常跳转URL,爬取quotes.toscrape.com仅能获取第一页如何解决?

Scrapy仅爬取第一页问题修复方案

核心问题及修复点

    1. allowed_domains配置不符合规范
      allowed_domains仅接收纯域名,不需要携带路径后缀,你配置中末尾的/会触发Scrapy的跨域拦截规则,所有分页请求会被直接过滤,不会发出。
      修复后配置:allowed_domains = ['quotes.toscrape.com']
    1. 分页请求代码缩进错误
      你把下一页请求的逻辑写在了遍历单条quote的for循环内部,会导致两个问题:一是每处理一条quote就重复发起一次下一页请求,产生大量无效请求;二是如果页面quote为空,永远不会触发分页请求。
      修复方法:将下一页相关代码整体移出for循环,取消一层缩进。
    1. 缺少下一页链接判空逻辑
      最后一页不存在下一页按钮时,next_page_url会返回None,直接调用urljoin会抛出异常,需要增加非空判断再发起请求。

修复后完整代码

import scrapy


class QuotesSpider(scrapy.Spider):
    name = 'quotes'
    allowed_domains = ['quotes.toscrape.com']
    start_urls = ['http://quotes.toscrape.com/']

    def parse(self, response):
        quotes = response.xpath('//*[@class="quote"]')
        for quote in quotes :
            text = quote.xpath('.//*[@class="text"]/text()').extract_first()
            author = quote.xpath('.//*[@class="author"]/text()').extract_first()
            tags = quote.xpath('.//*[@class="keywords"]/@content').extract_first()
            yield{
                'text':text,
                'author':author,
                'tags':tags
            }
       
        next_page_url = response.xpath('//*[@class="next"]/a/@href').extract_first()
        if next_page_url:
            absolute_next_page_url = response.urljoin(next_page_url)
            yield scrapy.Request(absolute_next_page_url)

修改完成后重新运行即可正常爬取全站10页的全部名言数据。

内容的提问来源于stack exchange,提问作者Faycal Faycal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 15:15:05