Scrapy Spider回调未触发:爬取GitHub仓库XML文件遇阻
问题分析与解决方案
核心问题1:重复请求被Scrapy自动过滤
你在parse_level中发起了针对当前页面URL的重复请求:
yield scrapy.Request(response.url, callback=self.parse_docs)
Scrapy默认开启重复请求过滤机制,相同URL的请求会被直接忽略,因此parse_docs回调永远不会触发。
核心问题2:爬取规则缺少follow=True
你的repo_rule和pagination_rule都未设置follow=True:
repo_rule提取仓库链接后,不会进入仓库页面,后续level_rule无法匹配到仓库内的/level1路径pagination_rule提取翻页链接后,不会跳转下一页,只能爬取第一页的仓库
修正后的完整代码
import scrapy from scrapy.crawler import CrawlerProcess from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor repo_rule = Rule( LinkExtractor( restrict_xpaths="//a[@itemprop='name codeRepository']", restrict_text=r"ELTeC-.+" ), follow=True # 开启跟进,进入仓库页面 ) pagination_rule = Rule( LinkExtractor(restrict_xpaths="//a[@class='next_page']"), follow=True # 开启跟进,翻页爬取所有仓库 ) level_rule = Rule( LinkExtractor(allow=r"/level1"), follow=False, # level1页面无需继续跟进,直接处理 callback="parse_level" ) class ELTecSpider(CrawlSpider): """Scrapy CrawlSpider for crawling the ELTec repo.""" name = "eltec" start_urls = ["https://github.com/orgs/COST-ELTeC/repositories"] rules = [repo_rule, pagination_rule, level_rule] def parse_level(self, response): # 直接用当前响应处理,无需重复发请求 yield from self.parse_docs(response) def parse_docs(self, response): # 提取level1目录下的所有XML文件链接 xml_links = response.xpath( "//a[@class='Link--primary'][contains(@href, '.xml')]/@href" ).getall() for link in xml_links: full_url = response.urljoin(link) print("找到XML文件:", full_url) # 发起请求爬取XML内容 yield scrapy.Request(full_url, callback=self.parse_xml_content) def parse_xml_content(self, response): # 解析XML内容,示例提取根节点信息(可根据实际XML结构调整) root = response.xpath('/*') item = { "xml_url": response.url, "root_tag": root.xpath('name(.)').get(), # 可添加更多XML字段提取逻辑 } yield item process = CrawlerProcess( settings={ "FEEDS": { "eltec_xmls.json": { "format": "json", "overwrite": True }, }, # 可选:添加请求延迟,避免触发GitHub反爬 "DOWNLOAD_DELAY": 1, } ) process.crawl(ELTecSpider) process.start()
关键修改说明
- 移除重复请求,直接在
parse_level中调用parse_docs处理当前响应 - 给
repo_rule和pagination_rule添加follow=True,确保能进入仓库页面并翻页 - 优化XML链接提取逻辑,直接筛选
.xml后缀的链接 - 添加
parse_xml_content方法处理XML文件内容,可根据实际XML结构调整提取规则
内容的提问来源于stack exchange,提问作者lupl
相关产品推荐
相关产品推荐

