Scrapy CrawlProcess的crawl()函数偶发无限挂起问题排查求助
Scrapy结合Selenium下载器爬取时crawl()无限挂起的调试与超时设置
问题背景
使用CrawlProcess搭配Selenium下载器运行Scrapy爬虫,多数场景下运行正常,但偶尔调用crawl()后会无限挂起,且无任何异常抛出。尝试通过spider_idle信号排查未生效,最终定位到代码卡在start_requests函数中,连初始打印语句都未执行。
更新1:spider_idle信号未触发
配置了spider_idle信号但未触发,代码如下:
@classmethod def from_crawler(cls, crawler, *args, **kwargs): print("-- Initiating signals from_crawler function") spider = super(MySpider, cls).from_crawler(crawler, *args, **kwargs) crawler.signals.connect(spider.handleSpiderIdle, signal=signals.spider_idle) return spider
对应的信号处理函数未被调用:
def handleSpiderIdle(self, spider): '''Handle spider idle event.''' print(f'\nSpider idle: {spider.name}. Restarting it... ')
更新2:定位到卡在start_requests函数
排查发现代码卡在start_requests阶段,初始打印语句无输出:
def start_requests(self): print("-- Initial Request Started --") # 未执行此打印 yield SeleniumRequest( url=self.start_urls[0], callback=self.parse, wait_time=2, ) print("-- Initial Request Passed --")
调试方法
1. 开启Scrapy详细日志
启动CrawlProcess前配置日志级别为DEBUG,捕获底层运行细节:
from scrapy.utils.log import configure_logging import logging configure_logging(install_root_handler=False) logging.basicConfig( format='%(levelname)s: %(message)s', level=logging.DEBUG )
重点关注Selenium下载器的初始化、浏览器启动、请求发送阶段的日志,排查是否存在资源阻塞或隐性超时问题。
2. 检查Selenium资源管理
- 强制销毁浏览器实例:避免未关闭的浏览器进程占用资源,导致后续请求无法启动新进程
def closed(self, reason): if hasattr(self, 'driver'): self.driver.quit() - 限制并发请求数:Selenium属于资源密集型操作,单线程运行可减少资源竞争
在settings.py中设置:CONCURRENT_REQUESTS = 1 DOWNLOAD_DELAY = 3
3. 为start_requests添加异常捕获
直接在start_requests中包裹异常处理,捕获隐性错误:
def start_requests(self): try: print("-- Initial Request Started --") yield SeleniumRequest( url=self.start_urls[0], callback=self.parse, wait_time=2, ) print("-- Initial Request Passed --") except Exception as e: print(f"start_requests错误: {str(e)}") logging.error(f"start_requests错误: {str(e)}", exc_info=True) raise e
为crawl()设置超时
Scrapy的CrawlProcess无直接超时参数,可通过Python线程/进程实现超时控制:
方法1:使用threading实现超时
import threading from scrapy.crawler import CrawlerProcess from scrapy.utils.engine import get_engine from myspiders import MySpider def run_spider(): process = CrawlerProcess(settings={ # 你的爬虫配置 }) process.crawl(MySpider) process.start() # 设置10分钟超时 spider_thread = threading.Thread(target=run_spider) spider_thread.start() spider_thread.join(timeout=600) if spider_thread.is_alive(): print("爬虫超时,强制终止") engine = get_engine() engine.stop()
方法2:使用multiprocessing实现超时
import multiprocessing from scrapy.crawler import CrawlerProcess from myspiders import MySpider def run_spider(): process = CrawlerProcess(settings={ # 你的爬虫配置 }) process.crawl(MySpider) process.start() spider_process = multiprocessing.Process(target=run_spider) spider_process.start() spider_process.join(timeout=600) if spider_process.is_alive(): print("爬虫超时,强制终止") spider_process.terminate() spider_process.join()
内容的提问来源于stack exchange,提问作者Arman Avetisyan
相关产品推荐
相关产品推荐

