You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django+Celery+Scrapy架构下ReactorNotRestartable错误求助

Fixing twisted.internet.error.ReactorNotRestartable in Your Django-Celery-Scrapy Pipeline

The ReactorNotRestartable error occurs because Twisted's reactor (the core event loop Scrapy relies on) can only be started once per process. Since Celery workers are long-lived processes that handle multiple tasks, if your first task runs the spider and stops the reactor, any subsequent task trying to restart it in the same worker will fail with this error.

Here are two reliable solutions to fix this:


Solution 1: Run the Scrapy Spider as a Subprocess

The simplest fix is to invoke the scrapy crawl command directly via a subprocess. This isolates each spider run in its own process, so the reactor doesn't interfere with future Celery tasks.

Update your tasks.py like this:

from celery import shared_task
import subprocess
from pathlib import Path

@shared_task
def positions(keyword_chunk):
    # Adjust this path to point to your Scrapy project root
    scrapy_project_root = Path(__file__).parent.parent

    for keyword in keyword_chunk:
        try:
            # Execute the spider via the scrapy command line tool
            result = subprocess.run(
                [
                    "scrapy",
                    "crawl",
                    "MySpider",
                    "-a", f"keyword={keyword}"  # Pass keyword to your spider
                ],
                cwd=scrapy_project_root,
                check=True,
                capture_output=True,
                text=True
            )
            # Optional: Log successful runs
            print(f"Successfully scraped keyword: {keyword}\nOutput: {result.stdout}")
        except subprocess.CalledProcessError as e:
            print(f"Failed to scrape keyword {keyword}: {e.stderr}")
            raise  # Let Celery handle task failure tracking

Solution 2: Spawn a Separate Process for Each Spider Run

If you prefer using Scrapy's programmatic API, use Python's multiprocessing module to run each spider in an isolated child process. This keeps the Celery worker's reactor untouched.

Update your tasks.py like this:

from celery import shared_task
from multiprocessing import Process
from scrapy.crawler import CrawlerProcess
from scrapy.utils.project import get_project_settings
from tracker.spiders.myspider import MySpider  # Import your spider class

def run_spider(keyword):
    settings = get_project_settings()
    process = CrawlerProcess(settings)
    process.crawl(MySpider, keyword=keyword)
    process.start()  # Blocks until the spider finishes

@shared_task
def positions(keyword_chunk):
    processes = []
    for keyword in keyword_chunk:
        p = Process(target=run_spider, args=(keyword,))
        p.start()
        processes.append(p)
    
    # Wait for all spiders to complete before finishing the task
    for p in processes:
        p.join()

Key Notes

  • Ensure your Celery worker has access to the Scrapy environment (the scrapy command should be in its PATH, and project settings are correctly configured).
  • Avoid running the spider directly in the Celery worker process with CrawlerProcess.start()—this stops the reactor and breaks future tasks in that worker instance.
  • Both solutions prevent reactor reuse issues by isolating each spider run from the long-lived Celery worker process.

内容的提问来源于stack exchange,提问作者Роман Денисенко

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:29:46