You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Spider回调未触发:爬取GitHub仓库XML文件遇阻

问题分析与解决方案

核心问题1:重复请求被Scrapy自动过滤

你在parse_level中发起了针对当前页面URL的重复请求:

yield scrapy.Request(response.url, callback=self.parse_docs)

Scrapy默认开启重复请求过滤机制,相同URL的请求会被直接忽略,因此parse_docs回调永远不会触发。

核心问题2:爬取规则缺少follow=True

你的repo_rule和pagination_rule都未设置follow=True:

  • repo_rule提取仓库链接后,不会进入仓库页面,后续level_rule无法匹配到仓库内的/level1路径
  • pagination_rule提取翻页链接后,不会跳转下一页,只能爬取第一页的仓库

修正后的完整代码

import scrapy
from scrapy.crawler import CrawlerProcess
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor


repo_rule = Rule(
    LinkExtractor(
        restrict_xpaths="//a[@itemprop='name codeRepository']",
        restrict_text=r"ELTeC-.+"
    ),
    follow=True  # 开启跟进,进入仓库页面
)

pagination_rule = Rule(
    LinkExtractor(restrict_xpaths="//a[@class='next_page']"),
    follow=True  # 开启跟进,翻页爬取所有仓库
)

level_rule = Rule(
    LinkExtractor(allow=r"/level1"),
    follow=False,  # level1页面无需继续跟进,直接处理
    callback="parse_level"
)


class ELTecSpider(CrawlSpider):
    """Scrapy CrawlSpider for crawling the ELTec repo."""

    name = "eltec"
    start_urls = ["https://github.com/orgs/COST-ELTeC/repositories"]
    rules = [repo_rule, pagination_rule, level_rule]

    def parse_level(self, response):
        # 直接用当前响应处理,无需重复发请求
        yield from self.parse_docs(response)

    def parse_docs(self, response):
        # 提取level1目录下的所有XML文件链接
        xml_links = response.xpath(
            "//a[@class='Link--primary'][contains(@href, '.xml')]/@href"
        ).getall()
        
        for link in xml_links:
            full_url = response.urljoin(link)
            print("找到XML文件:", full_url)
            # 发起请求爬取XML内容
            yield scrapy.Request(full_url, callback=self.parse_xml_content)

    def parse_xml_content(self, response):
        # 解析XML内容,示例提取根节点信息(可根据实际XML结构调整)
        root = response.xpath('/*')
        item = {
            "xml_url": response.url,
            "root_tag": root.xpath('name(.)').get(),
            # 可添加更多XML字段提取逻辑
        }
        yield item


process = CrawlerProcess(
    settings={
        "FEEDS": {
            "eltec_xmls.json": {
                "format": "json",
                "overwrite": True
            },
        },
        # 可选:添加请求延迟,避免触发GitHub反爬
        "DOWNLOAD_DELAY": 1,
    }
)

process.crawl(ELTecSpider)
process.start()

关键修改说明

  1. 移除重复请求,直接在parse_level中调用parse_docs处理当前响应
  2. 给repo_rule和pagination_rule添加follow=True,确保能进入仓库页面并翻页
  3. 优化XML链接提取逻辑,直接筛选.xml后缀的链接
  4. 添加parse_xml_content方法处理XML文件内容,可根据实际XML结构调整提取规则

内容的提问来源于stack exchange,提问作者lupl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 07:34:53