You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多域名子站爬取问题及爬虫重启方案求助

解决Scrapy多站点爬取子URL的问题

Hey there! Let's break down your problems and fix them step by step — you're already on the right track, we just need to tweak a few key parts of your code.

核心问题分析

你目前碰到两个关键卡点:

  1. 爬虫对部分传入链接不生效,日志显示爬取了站点但没正确提取所有子链接
  2. 无法通过循环逐个处理站点,实现“每个站点爬完后独立获取子URL列表”的需求

这些问题主要源于全局变量误用和Scrapy爬虫实例的配置逻辑没匹配多站点场景,咱们一一解决。

优化后的完整代码

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.crawler import CrawlerProcess
import re

# 用字典存储每个站点的子URL,键为域名,值为去重后的链接集合
site_suburls = {}

def extract_base_domain(url):
    """统一提取URL的基础域名,用于配置allowed_domains"""
    cleaned_url = re.sub(r'/$', '', url)
    cleaned_url = re.sub(r'^https?://', '', cleaned_url)
    cleaned_url = re.sub(r'^www\.', '', cleaned_url)
    return cleaned_url

class SiteCrawler(CrawlSpider):
    name = "site_crawler"
    
    def __init__(self, start_url, *args, **kwargs):
        super().__init__(*args, **kwargs)
        # 为当前爬虫实例单独配置起始URL和目标域名
        self.start_urls = [start_url]
        self.base_domain = extract_base_domain(start_url)
        self.allowed_domains = [self.base_domain]
        
        # 只提取当前域名下的链接,自动去重
        self.link_extractor = LinkExtractor(allow_domains=self.allowed_domains, unique=True)
        
        # 配置爬虫规则:提取链接后递归跟进爬取,处理每个页面
        self.rules = [
            Rule(self.link_extractor, callback='parse_subsite', follow=True)
        ]
        
        # 初始化当前站点的链接存储集合
        site_suburls[self.base_domain] = set()
        
        # 必须重新编译规则(Scrapy要求修改rules后执行此操作)
        self._compile_rules()

    def parse_subsite(self, response):
        # 将当前页面的URL添加到对应站点的集合中
        site_suburls[self.base_domain].add(response.url)

if __name__ == "__main__":
    # 要爬取的目标站点列表
    target_sites = ['https://bernd-lange.de/', 'https://markus-pieper.eu/']
    
    process = CrawlerProcess()
    
    # 为每个站点创建独立的爬虫实例
    for site_url in target_sites:
        process.crawl(SiteCrawler, start_url=site_url)
    
    # 启动所有爬虫,阻塞直到全部完成
    process.start()
    
    # 输出最终爬取结果
    print("=== 爬取结果汇总 ===")
    for domain, urls in site_suburls.items():
        print(f"\n域名 {domain} 的子URL共 {len(urls)} 个:")
        for url in sorted(urls):
            print(f"- {url}")

关键修改说明

  1. 独立爬虫实例:每个站点对应一个专属的SiteCrawler实例,在__init__中动态配置start_urls和allowed_domains,确保爬虫只专注于当前目标站点,不会混淆其他站点的链接
  2. 结果分离存储:用字典site_suburls替代全局列表,每个域名对应自己的链接集合,方便后续单独处理每个站点的结果
  3. 修复爬虫规则:设置follow=True让爬虫递归爬取子页面的链接,确保能获取到所有子站URL;同时调用_compile_rules()重新编译规则,适配动态修改的rules配置
  4. 统一域名提取:封装extract_base_domain函数,避免重复代码,确保域名提取逻辑一致

额外实用提示

  • 如果需要限制爬取深度,可以在Rule中添加depth_limit参数,比如Rule(..., follow=True, depth_limit=2)表示只爬取2层深度的页面
  • 为避免被网站封禁,建议在配置中开启ROBOTSTXT_OBEY = True(默认开启),并添加DOWNLOAD_DELAY = 2让爬虫每次请求后等待2秒
  • 如果部分站点爬取异常,可以检查网站的robots协议,或者在LinkExtractor中调整allow/deny规则过滤不需要的链接

内容的提问来源于stack exchange,提问作者Express

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:38:01