You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Python:重构parse方法后爬虫停止工作的原因

重构Scrapy爬虫后无法跟进链接的问题分析

原正常运行代码

class MySpider(BaseScrapper):
    name = "my_spider"

    def parse(self, response, **kwargs):
        self.logger.info(f"Parse: Processing {response.url}")

        yield ScrapyItem(
            source=response.meta["source"],
            url=response.url,
            html=response.text,
        )

        links = self.extract_links(response)

        self.logger.info(f"Extracted {len(links)} links from {response.url}")

        for link in links:
            self.logger.info(f"Following link: {link.url}")
            yield scrapy.Request(
                url=link.url,
                callback=self.parse,
                meta={
                    "source": response.meta["source"],
                },
            )

重构后异常代码

def follow_links(self, response, links):
        self.logger.info(f"Following {len(links)} links from {response.url}")
        for link in links:
            self.logger.info(f"Following link: {link.url}")
            yield scrapy.Request(
                url=link,
                callback=self.parse,
                meta={
                    "source": response.meta["source"],
                },
            )

    def extract_and_follow_links(self, response):
        links = self.extract_links(response)
        self.logger.info(f"Extracted {len(links)} links from {response.url}")

        # TODO: Save the links to the database
        self.follow_links(response, links)

    def parse(self, response, **kwargs):
        self.logger.info(f"Parse: Processing {response.url}")

        yield ScrapyItem(
            source=response.meta["source"],
            url=response.url,
            html=response.text,
        )
    
        self.extract_and_follow_links(response)

问题原因

  1. 生成器结果未传递:follow_links是生成器函数,直接调用它只会创建生成器对象,不会自动将内部生成的scrapy.Request传递给Scrapy引擎。原代码中parse函数直接yield Request,而重构后extract_and_follow_links仅调用生成器却未传递结果,导致新链接的请求从未被调度。
  2. URL参数错误:follow_links里的url=link应为url=link.url,原代码使用link.url,重构后误写为link,即使修复生成器问题,也会因URL格式错误导致请求失败。

修复方案

修正后的代码如下:

def follow_links(self, response, links):
        self.logger.info(f"Following {len(links)} links from {response.url}")
        for link in links:
            self.logger.info(f"Following link: {link.url}")
            yield scrapy.Request(
                url=link.url,  # 修正URL参数
                callback=self.parse,
                meta={
                    "source": response.meta["source"],
                },
            )

    def extract_and_follow_links(self, response):
        links = self.extract_links(response)
        self.logger.info(f"Extracted {len(links)} links from {response.url}")

        # TODO: Save the links to the database
        yield from self.follow_links(response, links)  # 传递生成器结果

    def parse(self, response, **kwargs):
        self.logger.info(f"Parse: Processing {response.url}")

        yield ScrapyItem(
            source=response.meta["source"],
            url=response.url,
            html=response.text,
        )
    
        yield from self.extract_and_follow_links(response)  # 将结果传递到parse的输出流

关键修复点说明

  • 使用yield from替代直接调用生成器函数,确保生成器中的scrapy.Request被传递给Scrapy引擎,新链接的请求才会被调度执行。
  • 修正url=link为url=link.url,保证请求使用正确的URL格式。

内容的提问来源于stack exchange,提问作者inquilabee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 07:35:39