Scrapy中CrawlSpider自定义parse函数未被调用问题求助
问题排查与修复方案
核心原因:CrawlSpider不允许重写parse方法
Scrapy的CrawlSpider类内置了默认的parse实现,用于处理用户定义的Rule规则匹配链接,如果你自定义了parse方法会覆盖原有逻辑,直接导致爬虫流程中断。你当前代码没有定义任何Rule规则,完全不需要继承CrawlSpider,直接继承基础的scrapy.Spider类即可解决parse不触发的问题。
其他需要修正的代码问题
- 语法错误:
if URL.split('/')[2] == "www.ricetta.it":后无缩进执行块,且你的起始域名是ricetta.it不是www.ricetta.it,原有判断条件永远不成立 - 属性调用错误:
response.request.url是属性不是方法,不能加()调用 - 提取逻辑错误:
recipes = response.xpath('xxx').extract()已经得到字符串列表,后续不能再调用.extract(),也不能直接用字符串拼接xpath表达式,应该循环selector对象而非提前提取的字符串 - 翻页逻辑位置错误:翻页请求不应该放在循环每一条菜谱的逻辑里,会重复生成大量无效重复请求
修复后可运行代码
import scrapy class SpiderRecipe(scrapy.Spider): name = "recipe" start_urls = [ 'https://ricetta.it/dolci?page=1', ] def parse(self,response): current_url = response.request.url if current_url.split('/')[2] == "ricetta.it": # 直接获取selector列表,不提前extract recipes = response.xpath('//div[contains(@class,"row")]/div[contains(@class,"post-img-left")]') # 遍历selector对象 for recipe in recipes: title = recipe.xpath('.//a[contains(@class, "post-title")]/text()').get() image = recipe.xpath('.//div[contains(@class,"videoContainer")]/img/@src').get() description = recipe.xpath('.//p[contains(@class,"post-excerpt")]/text()').get() yield { 'Title': title, 'Image': image, 'Description': description, } # 翻页逻辑放在循环外,避免重复生成请求 page = int(current_url.split('=')[1]) + 1 if page <= 148: yield scrapy.Request(f"{current_url.split('=')[0]}={page}", callback=self.parse, dont_filter=True)
启动方式
仍然执行原命令即可正常触发解析逻辑:scrapy runspider recipe.py -o output.json
内容的提问来源于stack exchange,提问作者Ele975
相关产品推荐
相关产品推荐

