You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy提取网页链接时仅处理首个匹配链接的问题排查

问题原因与解决办法
  • 核心原因:Scrapy的CrawlSpider规则里,一旦给Rule指定了callback参数,follow参数默认会设为False。这就导致爬虫爬完第一个匹配category/interviews的页面后,不会从该页面里继续提取符合规则的链接,直接停止了爬取。而如果不指定callback,follow默认是True,所以能正常遍历所有匹配规则的页面。

  • 解决办法:在你的Rule里显式加上follow=True,这样即使指定了回调函数,爬虫也会继续从回调处理的页面中提取符合规则的链接,实现全量爬取。

修改后的代码如下:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.http import Request


class MagazineCrawler(CrawlSpider):
    name = "MagazineCrawler"
    allowed_domains = ["eu-startups.com"]
    start_urls = ["https://www.eu-startups.com"]

    rules = (
        Rule(LinkExtractor(allow=["category/interviews"]), callback="parse_category", follow=True),
    )

    def parse_category(self, response):
        xpath_links = "//div[@class='td_block_inner tdb-block-inner td-fix-index']//a[@class='td-image-wrap ']/@href"
        subpage_links = response.xpath(xpath_links).extract()

        # Follow each subpage link and yield requests to crawl them
        for link in subpage_links:
            yield Request(link)

内容的提问来源于stack exchange,提问作者CoolMathematician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:03:16