如何使用Scrapy仅获取站点地图URL而不爬取目标网页?
如何让Scrapy仅提取站点地图中的URL而不爬取网站页面
你的问题出在SitemapSpider的默认行为上:它会自动解析站点地图里的URL,然后发起请求爬取这些URL对应的页面,所以你的parse方法会收到这些页面的响应,导致输出了网站本身的URL。
解决方法
重写parse_sitemap方法,直接从站点地图条目里提取URL,不生成对这些URL的请求,就能避免触发页面爬取。
修改后的代码
import scrapy from scrapy.crawler import CrawlerProcess class SitemapSpider(scrapy.spiders.SitemapSpider): name = "sitemap_spider" def __init__(self, sitemap_url, *args, **kwargs): super(SitemapSpider, self).__init__(*args, **kwargs) self.sitemap_urls = [sitemap_url] self.extracted_urls = [] def parse_sitemap(self, response): # 自动解析站点地图(包括索引文件中的子地图),遍历所有条目 for item in self._parse_sitemap(response): url = item['loc'] print(url) self.extracted_urls.append(url) # 不yield任何请求,避免触发页面爬取 def run_sitemap_scraper(sitemap_url): process = CrawlerProcess() process.crawl(SitemapSpider, sitemap_url=sitemap_url) process.start() # 示例调用 run_sitemap_scraper("https://ferienparkguide.de/sitemap_index.xml")
关键说明
parse_sitemap是SitemapSpider专门处理站点地图响应的方法,默认逻辑是生成对条目中URL的请求,我们重写后改变了这个行为。self._parse_sitemap(response)会自动处理站点地图索引文件,解析所有子站点地图的条目,返回的每个条目字典里,loc字段就是站点地图中记录的URL。- 这里只提取并保存URL,不输出任何请求,Scrapy就不会去爬取这些URL对应的页面,只会输出站点地图中的所有目标URL。
内容的提问来源于stack exchange,提问作者hal1988
相关产品推荐
相关产品推荐

