Scrapy问题:日志显示已爬14091页但速率0页/分钟,CPU占满无爬取
Let's figure out why your Scrapy crawler is spiking to 100% CPU and grinding to a halt—those stats and log clues tell us exactly where to look.
First, let's connect the dots: You've got hundreds of connection timeouts/refusals, and the crawler's stuck at 0 pages per minute. That means your spider is wasting all its CPU retrying failed requests instead of crawling new pages, and the target site is probably throttling or blocking you. Here's how to fix it:
1. Dial Back Concurrent Requests & Add Delays
Your crawler's likely flooding the target site with too many requests at once, triggering anti-measures. Tweak these in your settings.py:
- Lower concurrent limits to reduce load on both your machine and the target:
CONCURRENT_REQUESTS = 8 CONCURRENT_REQUESTS_PER_DOMAIN = 4 - Add a download delay with randomization to act more like a human:
DOWNLOAD_DELAY = 2 RANDOMIZE_DOWNLOAD_DELAY = True
2. Stop Retrying Doomed Requests
332 timeouts and 146 connection refusals? Scrapy's wasting cycles retrying requests that will never succeed. Fix your retry settings:
- Cut down retry attempts:
RETRY_TIMES = 2 - Only retry recoverable HTTP errors (skip connection failures entirely):
RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429] - For extra control, write a custom downloader middleware that drops
ConnectErrororConnectionRefusedErrorrequests immediately—no retries at all.
3. Check for Queue Clogs or Memory Leaks
A stuck crawler could mean your request queue is backed up, or your spider code has a memory leak:
- Use Scrapy's stats to compare
scheduler/enqueuedvsscheduler/dequeued. If enqueued keeps growing but dequeued stops, your spider might be generating more requests than it can process, or getting stuck parsing certain pages. - Audit your
parsefunction: Are there infinite loops? Heavy regex or data processing that's hogging CPU? Offload big tasks to background threads if needed. - Keep an eye on memory usage—if it climbs nonstop, look for unclosed files, cached data that's never cleared, or global variables holding large datasets.
4. Beat Anti-Crawling Measures
The flood of connection errors suggests the target site has blocked your IP or is throttling you. Try these:
- Rotate user agents to avoid being flagged—use Scrapy's built-in middleware or a library to randomize your User-Agent header.
- Add proxy rotation if you're hitting IP bans—use a pool of proxies to spread out your requests.
- Check if the site requires cookies or session handling; mimic a real user's session flow instead of sending stateless requests.
Quick Debugging Tips
- Use
toporhtopwhile crawling to confirm the CPU is being eaten by your Scrapy process (not other system tasks). - Test a single request with
scrapy shellto see if the target site responds normally—if it's slow or refuses, that's a clear anti-crawl sign. - Crank up logging to
DEBUGlevel to see exactly where the crawler is getting stuck.
These changes should get your CPU usage back in line and your crawler moving again.
内容的提问来源于stack exchange,提问作者user9320130

