You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy问题:日志显示已爬14091页但速率0页/分钟,CPU占满无爬取

Let's figure out why your Scrapy crawler is spiking to 100% CPU and grinding to a halt—those stats and log clues tell us exactly where to look.

First, let's connect the dots: You've got hundreds of connection timeouts/refusals, and the crawler's stuck at 0 pages per minute. That means your spider is wasting all its CPU retrying failed requests instead of crawling new pages, and the target site is probably throttling or blocking you. Here's how to fix it:

1. Dial Back Concurrent Requests & Add Delays

Your crawler's likely flooding the target site with too many requests at once, triggering anti-measures. Tweak these in your settings.py:

  • Lower concurrent limits to reduce load on both your machine and the target:
    CONCURRENT_REQUESTS = 8
    CONCURRENT_REQUESTS_PER_DOMAIN = 4
    
  • Add a download delay with randomization to act more like a human:
    DOWNLOAD_DELAY = 2
    RANDOMIZE_DOWNLOAD_DELAY = True
    

2. Stop Retrying Doomed Requests

332 timeouts and 146 connection refusals? Scrapy's wasting cycles retrying requests that will never succeed. Fix your retry settings:

  • Cut down retry attempts:
    RETRY_TIMES = 2
    
  • Only retry recoverable HTTP errors (skip connection failures entirely):
    RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429]
    
  • For extra control, write a custom downloader middleware that drops ConnectError or ConnectionRefusedError requests immediately—no retries at all.

3. Check for Queue Clogs or Memory Leaks

A stuck crawler could mean your request queue is backed up, or your spider code has a memory leak:

  • Use Scrapy's stats to compare scheduler/enqueued vs scheduler/dequeued. If enqueued keeps growing but dequeued stops, your spider might be generating more requests than it can process, or getting stuck parsing certain pages.
  • Audit your parse function: Are there infinite loops? Heavy regex or data processing that's hogging CPU? Offload big tasks to background threads if needed.
  • Keep an eye on memory usage—if it climbs nonstop, look for unclosed files, cached data that's never cleared, or global variables holding large datasets.

4. Beat Anti-Crawling Measures

The flood of connection errors suggests the target site has blocked your IP or is throttling you. Try these:

  • Rotate user agents to avoid being flagged—use Scrapy's built-in middleware or a library to randomize your User-Agent header.
  • Add proxy rotation if you're hitting IP bans—use a pool of proxies to spread out your requests.
  • Check if the site requires cookies or session handling; mimic a real user's session flow instead of sending stateless requests.

Quick Debugging Tips

  • Use top or htop while crawling to confirm the CPU is being eaten by your Scrapy process (not other system tasks).
  • Test a single request with scrapy shell to see if the target site responds normally—if it's slow or refuses, that's a clear anti-crawl sign.
  • Crank up logging to DEBUG level to see exactly where the crawler is getting stuck.

These changes should get your CPU usage back in line and your crawler moving again.

内容的提问来源于stack exchange,提问作者user9320130

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:38:51