You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy中CrawlSpider自定义parse函数未被调用问题求助

问题排查与修复方案

核心原因:CrawlSpider不允许重写parse方法

Scrapy的CrawlSpider类内置了默认的parse实现,用于处理用户定义的Rule规则匹配链接,如果你自定义了parse方法会覆盖原有逻辑,直接导致爬虫流程中断。你当前代码没有定义任何Rule规则,完全不需要继承CrawlSpider,直接继承基础的scrapy.Spider类即可解决parse不触发的问题。

其他需要修正的代码问题

  • 语法错误:if URL.split('/')[2] == "www.ricetta.it": 后无缩进执行块,且你的起始域名是ricetta.it不是www.ricetta.it,原有判断条件永远不成立
  • 属性调用错误:response.request.url是属性不是方法,不能加()调用
  • 提取逻辑错误:recipes = response.xpath('xxx').extract()已经得到字符串列表,后续不能再调用.extract(),也不能直接用字符串拼接xpath表达式,应该循环selector对象而非提前提取的字符串
  • 翻页逻辑位置错误:翻页请求不应该放在循环每一条菜谱的逻辑里,会重复生成大量无效重复请求

修复后可运行代码

import scrapy

class SpiderRecipe(scrapy.Spider):
    name = "recipe"
    start_urls = [
        'https://ricetta.it/dolci?page=1',
    ]

    def parse(self,response):
        current_url = response.request.url
        if current_url.split('/')[2] == "ricetta.it":
            # 直接获取selector列表,不提前extract
            recipes = response.xpath('//div[contains(@class,"row")]/div[contains(@class,"post-img-left")]')
            # 遍历selector对象
            for recipe in recipes:
                title = recipe.xpath('.//a[contains(@class, "post-title")]/text()').get()
                image = recipe.xpath('.//div[contains(@class,"videoContainer")]/img/@src').get()
                description = recipe.xpath('.//p[contains(@class,"post-excerpt")]/text()').get()
                yield {
                    'Title': title,
                    'Image': image,
                    'Description': description,
                }
        # 翻页逻辑放在循环外,避免重复生成请求
        page = int(current_url.split('=')[1]) + 1
        if page <= 148:
            yield scrapy.Request(f"{current_url.split('=')[0]}={page}", callback=self.parse, dont_filter=True)

启动方式

仍然执行原命令即可正常触发解析逻辑:
scrapy runspider recipe.py -o output.json

内容的提问来源于stack exchange,提问作者Ele975

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 04:24:06