Scrapy Cloud广域爬取速度骤降求助:10k域名/6单元/900并发
Hey there, let's dig into why your broad crawl is slowing down after the initial hour and how to get it back up to speed. Your setup has a few conflicting configurations and common broad crawl pain points that are almost certainly contributing to the slowdown.
Key Causes of the Slowdown
1. Autothrottle vs. Global Concurrency Misalignment
You’ve set CONCURRENT_REQUESTS=900 for global concurrency, but AUTOTHROTTLE_TARGET_CONCURRENCY=1 tells Scrapy to only maintain 1 concurrent request per domain. For a broad crawl across 10,000 domains, this means your 900 concurrent slots can only be fully utilized if you’re actively crawling 900+ domains at once. As the crawl progresses, many domains will be fully scraped, and remaining ones may trigger autothrottle’s adaptive delays (due to slow responses or implicit anti-scraping measures), reducing the number of active, high-throughput domains. The result? Most of your 900 concurrent slots sit idle, leading to that 50 items/minute crawl rate.
2. Depleting Domain Queue & Uneven URL Distribution
In the first hour, you’re hitting domains with large numbers of accessible URLs, so your concurrency runs hot. After 3 hours, many domains are already fully crawled, leaving you with domains that either have very few pages left to scrape or require deep, slow traversal. With fewer high-volume domains in the queue, your crawler can’t keep all 900 concurrent requests busy.
3. Widespread Anti-Scraping Measures
Broad crawls are a red flag for most websites. Over time, more and more domains will start limiting your requests:
- Returning
429 Too Many Requestsor503 Service Unavailablestatus codes - Adding intentional delays to responses
- Blocking your crawler’s IP entirely
Each of these scenarios leads to retries, idle connections, or wasted requests that drag down your overall throughput.
4. Scrapy Cloud Resource Constraints
6 running units might not be enough to handle 900 concurrent requests, especially if each request involves parsing metadata, storing results, and handling middleware. If your units are hitting CPU, memory, or network limits, requests will queue up internally, increasing latency and reducing effective throughput.
Optimization Tips to Boost Crawl Speed
1. Fix Autothrottle & Concurrency Settings
- Adjust per-domain concurrency: Raise
AUTOTHROTTLE_TARGET_CONCURRENCYto3-5(instead of 1). This lets you maintain a small number of concurrent requests per domain without triggering aggressive anti-scraping, while better utilizing your global 900 concurrent slots. - Tweak autothrottle bounds: Set
AUTOTHROTTLE_START_DELAY=1andAUTOTHROTTLE_MAX_DELAY=10to prevent excessive delays from slowing down responsive domains. - Use per-domain concurrency explicitly: If you disable autothrottle (not recommended for broad crawls), set
CONCURRENT_REQUESTS_PER_DOMAIN=3-5to balance speed and anti-scraping risk.
2. Optimize Queue & Crawl Strategy
- Prioritize high-volume domains: Use Scrapy’s
PriorityQueueto assign higher priority to domains with more pending URLs. This keeps your concurrency slots filled with active, productive domains. - Limit crawl depth if needed: If "full site crawl" doesn’t require deep nested pages, set
DEPTH_LIMIT=3-5to avoid wasting resources on low-value, deep pages that add minimal metadata. - Clean up invalid URLs: Enable middleware to filter out URLs that consistently return 404s or redirect loops, so your crawler doesn’t waste requests on dead ends.
3. Mitigate Anti-Scraping Blocks
- Use rotating proxies: Configure Scrapy Cloud to use a proxy pool to avoid IP bans across multiple domains. This is critical for broad crawls where a single IP can be flagged quickly.
- Randomize User-Agents: Add middleware to rotate User-Agent strings with each request, mimicking real browser traffic.
- Tune retry logic: Adjust
RETRY_TIMES=3andRETRY_HTTP_CODESto include429, 503, 504—use exponential backoff (RETRY_DELAY=1,RETRY_BACKOFF=2) to give domains time to unblock you before retrying.
4. Scale Scrapy Cloud Resources
- Increase running units: If 6 units are hitting resource limits, add more units to distribute the concurrent request load.
- Adjust global concurrency: If you’re seeing high latency or request timeouts, lower
CONCURRENT_REQUESTSto600—sometimes a slightly lower, stable concurrency rate delivers higher throughput than an overloaded max setting. - Check logs for bottlenecks: Enable
LOG_LEVEL='INFO'to track request status codes, retry counts, and response times. Look for domains that consistently trigger retries or slow responses, and consider adding them to aROBOTSTXT_OBEYexclusion list or reducing crawl rate for those specific sites.
内容的提问来源于stack exchange,提问作者romain-lavoix

