Scrapy爬虫报错:Twisted非对象弱引用问题求助
Hey there, let's break down why both your CrawlerRunner and CrawlerProcess approaches are failing with that Twisted error. From the traceback snippet you shared, it's clear the issue is tied to Twisted's deferred execution system—let's go through the most likely fixes step by step.
Your Error Log
2018-04-18 23:55:46 [twisted] CRITICAL: Traceback (most recent call last):
File "/home/flo/PycharmProjects/Tensortest/Car_Analysis/lib/python3.6/site-packages/twisted/internet/defer.py", line 1386, in _inlineCallbacks
result...
Possible Fixes
1. Fix Dependency Version Conflicts
Scrapy and Twisted have strict version compatibility rules. If you're running a newer Twisted version with an older Scrapy (or vice versa), it can break the deferred callback system.
- First, check your current versions:
pip show scrapy twisted - For your 2018-era setup (Python 3.6), try installing a known-compatible combination:
This pair was widely tested and stable back then.pip install scrapy==1.5.1 twisted==18.7.0
2. Ensure Correct Reactor Handling
Both CrawlerRunner and CrawlerProcess handle the Twisted reactor differently—mixing up their usage can cause crashes.
Correct CrawlerRunner Implementation
CrawlerRunner doesn't start the reactor automatically, so you need to explicitly run it:
from scrapy.crawler import CrawlerRunner from scrapy.utils.log import configure_logging from twisted.internet import reactor from myproject.spiders import MySpider # Import your spider configure_logging() runner = CrawlerRunner() runner.crawl(MySpider) # Stop reactor when all crawlers finish d = runner.join() d.addBoth(lambda _: reactor.stop()) reactor.run() # Start the reactor
Correct CrawlerProcess Implementation
CrawlerProcess manages the reactor for you, so don't call reactor.run() manually:
from scrapy.crawler import CrawlerProcess from scrapy.utils.project import get_project_settings from myproject.spiders import MySpider process = CrawlerProcess(get_project_settings()) process.crawl(MySpider) process.start() # This starts the reactor internally
Pro tip: Never mix manual reactor calls with CrawlerProcess—it'll almost certainly cause conflicts.
3. Debug Your Spider Code
Sometimes the issue isn't with the crawler setup, but with your spider's logic. A broken callback, missing field, or unhandled exception in your spider/pipeline/middleware can trigger Twisted's critical error.
- Simplify your spider temporarily: Remove pipelines, middlewares, and complex callbacks. Run just a basic spider that makes one request and prints a response.
- If the simplified spider works, add back components one by one to find which part is causing the crash.
4. Reset Your Virtual Environment
Corrupted virtual environments can lead to weird, hard-to-debug issues. Try starting fresh:
# Deactivate current env if active deactivate # Create a new virtual environment python3.6 -m venv fresh_scrapy_env # Activate it source fresh_scrapy_env/bin/activate # Linux/macOS # fresh_scrapy_env\Scripts\activate # Windows # Install only the necessary dependencies pip install scrapy==1.5.1 twisted==18.7.0
Then test your crawler in this clean environment.
内容的提问来源于stack exchange,提问作者Fscir

