You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy仅获取站点地图URL而不爬取目标网页?

如何让Scrapy仅提取站点地图中的URL而不爬取网站页面

你的问题出在SitemapSpider的默认行为上:它会自动解析站点地图里的URL,然后发起请求爬取这些URL对应的页面,所以你的parse方法会收到这些页面的响应,导致输出了网站本身的URL。

解决方法

重写parse_sitemap方法,直接从站点地图条目里提取URL,不生成对这些URL的请求,就能避免触发页面爬取。

修改后的代码

import scrapy
from scrapy.crawler import CrawlerProcess

class SitemapSpider(scrapy.spiders.SitemapSpider):
    name = "sitemap_spider"
    
    def __init__(self, sitemap_url, *args, **kwargs):
        super(SitemapSpider, self).__init__(*args, **kwargs)
        self.sitemap_urls = [sitemap_url]
        self.extracted_urls = []

    def parse_sitemap(self, response):
        # 自动解析站点地图(包括索引文件中的子地图),遍历所有条目
        for item in self._parse_sitemap(response):
            url = item['loc']
            print(url)
            self.extracted_urls.append(url)
            # 不yield任何请求,避免触发页面爬取

def run_sitemap_scraper(sitemap_url):
    process = CrawlerProcess()
    process.crawl(SitemapSpider, sitemap_url=sitemap_url)
    process.start()

# 示例调用
run_sitemap_scraper("https://ferienparkguide.de/sitemap_index.xml")

关键说明

  • parse_sitemap是SitemapSpider专门处理站点地图响应的方法,默认逻辑是生成对条目中URL的请求,我们重写后改变了这个行为。
  • self._parse_sitemap(response)会自动处理站点地图索引文件,解析所有子站点地图的条目,返回的每个条目字典里,loc字段就是站点地图中记录的URL。
  • 这里只提取并保存URL,不输出任何请求,Scrapy就不会去爬取这些URL对应的页面,只会输出站点地图中的所有目标URL。

内容的提问来源于stack exchange,提问作者hal1988

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 18:15:04