Scrapy脚本运行爬虫出现ReactorNotRestartable错误求助
解决Scrapy运行脚本时的
ReactorNotRestartable错误 这个错误我之前也碰到过,本质是Twisted的反应器(Reactor)机制导致的——Twisted的Reactor是单实例设计,一旦停止就无法重新启动。咱们先拆解你的代码问题,再给你对应解决方案:
问题根源
你的代码里process.start()会一直阻塞到爬虫完全结束,此时Scrapy已经自动停止了Reactor。之后调用的process.stop()不仅多余,而且如果后续你尝试再次启动爬虫(比如测试场景下重复运行),就会触发ReactorNotRestartable错误。
解决方案分三种场景:
场景1:只需要运行一次爬虫
直接去掉多余的process.stop()即可,因为process.start()会在爬虫完成后自动停止Reactor:
from scrapy.crawler import CrawlerProcess from scrapy.utils.project import get_project_settings process = CrawlerProcess(get_project_settings()) process.crawl('followall', domain='scrapinghub.com') process.start() # 爬虫结束后会自动停止反应器,无需手动调用stop
场景2:需要在同一个脚本里运行多个爬虫
不要单独调用stop,一次性把所有爬虫添加到CrawlerProcess后再启动,Reactor会在所有爬虫都完成后才停止:
from scrapy.crawler import CrawlerProcess from scrapy.utils.project import get_project_settings process = CrawlerProcess(get_project_settings()) # 添加多个爬虫任务 process.crawl('followall', domain='scrapinghub.com') process.crawl('your_other_spider', domain='example.com') process.start() # 所有爬虫完成后自动停止反应器
场景3:需要重复运行爬虫(比如测试场景)
这种情况更推荐使用CrawlerRunner而非CrawlerProcess,因为它不会自动启动/停止Reactor,让你更灵活控制运行逻辑:
单次运行示例
from scrapy.crawler import CrawlerRunner from scrapy.utils.project import get_project_settings from twisted.internet import reactor runner = CrawlerRunner(get_project_settings()) # 运行爬虫,完成后触发停止反应器 d = runner.crawl('followall', domain='scrapinghub.com') d.addBoth(lambda _: reactor.stop()) reactor.run() # 手动启动反应器
重复多次运行示例
from scrapy.crawler import CrawlerRunner from scrapy.utils.project import get_project_settings from twisted.internet import reactor, defer runner = CrawlerRunner(get_project_settings()) @defer.inlineCallbacks def repeat_crawl(): # 这里可以自定义重复次数 for _ in range(3): yield runner.crawl('followall', domain='scrapinghub.com') reactor.stop() repeat_crawl() reactor.run()
内容的提问来源于stack exchange,提问作者Laxman
相关产品推荐
相关产品推荐

