You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy结合Django与huey任务队列运行时Reactor重启失败问题

问题描述

用Scrapy+Django搭建网络爬虫,将CrawlerRunner代码部署到huey任务队列后,本地运行正常,但服务器执行时出现异常:首次运行触发截断的Unhandled Error,第二次运行直接抛出ReactorNotRestartable异常。推测首次运行后Twisted Reactor未正常终止,导致10分钟后huey再次启动Reactor失败,怀疑多线程环境导致任务运行器与Twisted通信异常,寻求解决方案。

任务代码

from huey import crontab
from huey.contrib.djhuey import db_periodic_task, on_startup
from scrapy.crawler import CrawlerRunner
from scrapy.utils.log import configure_logging
from scrapy.utils.project import get_project_settings
from twisted.internet import reactor

from apps.core.tasks import CRONTAB_PERIODS
from apps.scrapers.crawler1 import Crawler1
from apps.scrapers.crawler2 import Crawler2
from apps.scrapers.crawler3 import Crawler3


@on_startup(name="scrape_all__on_startup")
@db_periodic_task(crontab(**CRONTAB_PERIODS["every_10_minutes"]))
def scrape_all():
    configure_logging()
    settings = get_project_settings()

    runner = CrawlerRunner(settings=settings)

    runner.crawl(Crawler1)
    runner.crawl(Crawler2)
    runner.crawl(Crawler3)

    defer = runner.join()
    defer.addBoth(lambda _: reactor.stop())

    reactor.run()

报错信息

首次运行报错(截断)

Unhandled Error
Traceback (most recent call last):
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/base.py", line 501, in fireEvent
    DeferredList(beforeResults).addCallback(self._continueFiring)
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/defer.py", line 532, in addCallback
    return self.addCallbacks(callback, callbackArgs=args, callbackKeywords=kwargs)
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/defer.py", line 512, in addCallbacks
    self._runCallbacks()
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/defer.py", line 892, in _runCallbacks
    current.result = callback(  # type: ignore[misc]
--- <exception caught here> ---
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/base.py", line 513, in _continueFiring
    callable(*args, **kwargs)
  File "/home/deployer/env/lib/python3.10/site-packages/twisted/internet/base.py", line 1314, in _reallyStartRunning
    self._handle...

第二次运行报错

ReactorNotRestartable: null
  File "huey/api.py", line 379, in _execute
    task_value = task.execute()
  File "huey/api.py", line 772, in execute
    return func(*args, **kwargs)
  File "huey/contrib/djhuey/__init__.py", line 135, in inner
    return fn(*args, **kwargs)
  File "apps/series/tasks.py", line 31, in scrape_all
    reactor.run()
  File "twisted/internet/base.py", line 1317, in run
    self.startRunning(installSignalHandlers=installSignalHandlers)
  File "twisted/internet/base.py", line 1299, in startRunning
    ReactorBase.startRunning(cast(ReactorBase, self))
  File "twisted/internet/base.py", line 843, in startRunning
    raise error.ReactorNotRestartable()

解决方案

核心问题分析

Twisted的Reactor是单实例且不可重启的,huey作为任务队列会在同一进程中重复执行scrape_all任务,首次运行后Reactor被停止,第二次调用reactor.run()必然触发异常。服务器环境的Unhandled Error大概率是信号处理、多线程环境与Reactor的兼容性冲突导致。

具体解决办法

1. 用CrawlerProcess替代CrawlerRunner(推荐)

CrawlerProcess会自动管理Reactor的启动和停止,无需手动调用reactor.run()和reactor.stop(),更适配定时任务场景:

from scrapy.crawler import CrawlerProcess  # 替换原CrawlerRunner导入

@on_startup(name="scrape_all__on_startup")
@db_periodic_task(crontab(**CRONTAB_PERIODS["every_10_minutes"]))
def scrape_all():
    configure_logging()
    settings = get_project_settings()

    process = CrawlerProcess(settings=settings)

    process.crawl(Crawler1)
    process.crawl(Crawler2)
    process.crawl(Crawler3)

    process.start()  # 自动启动Reactor,任务完成后自动停止

2. 子进程隔离爬虫逻辑(复杂场景适配)

如果必须使用CrawlerRunner,可通过子进程执行爬虫,彻底隔离Reactor实例:

import multiprocessing

def run_crawlers():
    configure_logging()
    settings = get_project_settings()
    runner = CrawlerRunner(settings=settings)
    runner.crawl(Crawler1)
    runner.crawl(Crawler2)
    runner.crawl(Crawler3)
    defer = runner.join()
    defer.addBoth(lambda _: reactor.stop())
    reactor.run()

@on_startup(name="scrape_all__on_startup")
@db_periodic_task(crontab(**CRONTAB_PERIODS["every_10_minutes"]))
def scrape_all():
    # 子进程运行爬虫,避免主进程Reactor终止后无法重启
    p = multiprocessing.Process(target=run_crawlers)
    p.start()
    p.join()

3. 禁用huey信号处理(辅助优化)

服务器环境的Unhandled Error可能源于huey信号处理与Twisted冲突,启动huey时添加参数禁用信号处理:

python manage.py run_huey --no-signal-handler

验证要点

  • 观察首次任务执行是否正常结束,无Unhandled Error
  • 等待第二次任务触发,确认不再出现ReactorNotRestartable异常
  • 检查爬虫数据是否正常入库

内容的提问来源于stack exchange,提问作者Ekin Ertaç

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 12:50:20