You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy CrawlProcess的crawl()函数偶发无限挂起问题排查求助

Scrapy结合Selenium下载器爬取时crawl()无限挂起的调试与超时设置

问题背景

使用CrawlProcess搭配Selenium下载器运行Scrapy爬虫,多数场景下运行正常,但偶尔调用crawl()后会无限挂起,且无任何异常抛出。尝试通过spider_idle信号排查未生效,最终定位到代码卡在start_requests函数中,连初始打印语句都未执行。

更新1:spider_idle信号未触发

配置了spider_idle信号但未触发,代码如下:

@classmethod
def from_crawler(cls, crawler, *args, **kwargs):
    print("-- Initiating signals from_crawler function")
    spider = super(MySpider, cls).from_crawler(crawler, *args, **kwargs)
    crawler.signals.connect(spider.handleSpiderIdle, signal=signals.spider_idle)
    return spider

对应的信号处理函数未被调用:

def handleSpiderIdle(self, spider):
   '''Handle spider idle event.'''  
   print(f'\nSpider idle: {spider.name}. Restarting it... ')

更新2:定位到卡在start_requests函数

排查发现代码卡在start_requests阶段,初始打印语句无输出:

def start_requests(self):
    print("-- Initial Request Started --") # 未执行此打印
    yield SeleniumRequest(
        url=self.start_urls[0],
        callback=self.parse,
        wait_time=2,
    )
    print("-- Initial Request Passed --")

调试方法

1. 开启Scrapy详细日志

启动CrawlProcess前配置日志级别为DEBUG,捕获底层运行细节:

from scrapy.utils.log import configure_logging
import logging

configure_logging(install_root_handler=False)
logging.basicConfig(
    format='%(levelname)s: %(message)s',
    level=logging.DEBUG
)

重点关注Selenium下载器的初始化、浏览器启动、请求发送阶段的日志,排查是否存在资源阻塞或隐性超时问题。

2. 检查Selenium资源管理

  • 强制销毁浏览器实例:避免未关闭的浏览器进程占用资源,导致后续请求无法启动新进程
    def closed(self, reason):
        if hasattr(self, 'driver'):
            self.driver.quit()
    
  • 限制并发请求数:Selenium属于资源密集型操作,单线程运行可减少资源竞争
    在settings.py中设置:
    CONCURRENT_REQUESTS = 1
    DOWNLOAD_DELAY = 3
    

3. 为start_requests添加异常捕获

直接在start_requests中包裹异常处理,捕获隐性错误:

def start_requests(self):
    try:
        print("-- Initial Request Started --")
        yield SeleniumRequest(
            url=self.start_urls[0],
            callback=self.parse,
            wait_time=2,
        )
        print("-- Initial Request Passed --")
    except Exception as e:
        print(f"start_requests错误: {str(e)}")
        logging.error(f"start_requests错误: {str(e)}", exc_info=True)
        raise e

为crawl()设置超时

Scrapy的CrawlProcess无直接超时参数,可通过Python线程/进程实现超时控制:

方法1:使用threading实现超时

import threading
from scrapy.crawler import CrawlerProcess
from scrapy.utils.engine import get_engine
from myspiders import MySpider

def run_spider():
    process = CrawlerProcess(settings={
        # 你的爬虫配置
    })
    process.crawl(MySpider)
    process.start()

# 设置10分钟超时
spider_thread = threading.Thread(target=run_spider)
spider_thread.start()
spider_thread.join(timeout=600)

if spider_thread.is_alive():
    print("爬虫超时,强制终止")
    engine = get_engine()
    engine.stop()

方法2:使用multiprocessing实现超时

import multiprocessing
from scrapy.crawler import CrawlerProcess
from myspiders import MySpider

def run_spider():
    process = CrawlerProcess(settings={
        # 你的爬虫配置
    })
    process.crawl(MySpider)
    process.start()

spider_process = multiprocessing.Process(target=run_spider)
spider_process.start()
spider_process.join(timeout=600)

if spider_process.is_alive():
    print("爬虫超时,强制终止")
    spider_process.terminate()
    spider_process.join()

内容的提问来源于stack exchange,提问作者Arman Avetisyan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 22:01:07