Scrapy-splash内存占用过高问题及重试请求延迟配置咨询
Let’s break this down step by step—first fixing the root memory issue (since 3 pages shouldn’t hog 10GB), then addressing the retry delay if you still need it as a fallback.
1. Better Solutions to Resolve the Memory Leak
Tune Splash Container Configuration
Splash is likely the source of the leak, even with low concurrent requests from Scrapy:
- Match Splash’s render slots to Scrapy’s concurrency: By default, Splash uses more slots than you might expect. When starting the container, add
--slots 3to align with yourCONCURRENT_REQUESTS=3setting. This prevents Splash from overloading itself with unneeded render processes:docker run -p 8050:8050 --slots 3 scrapinghub/splash - Set hard memory limits for Docker: Cap the container’s memory to avoid crashing your entire system. Use
--memoryand--memory-swapto trigger a clean OOM kill, plus--restart=alwaysto auto-restart the container without manual intervention:docker run -p 8050:8050 --slots 3 --memory=2g --memory-swap=2g --restart=always scrapinghub/splash - Limit render timeouts: Add
--max-timeout 30to prevent stuck renders from hanging around and consuming memory indefinitely.
Audit Custom Lua Scripts (If Used)
If you’re using a custom Lua script with Splash, make sure you’re cleaning up resources:
- Always call
splash:stop()after rendering a page to close the browser session and free memory. - Avoid storing large, unused objects in the script’s scope that won’t get garbage collected.
Optimize Scrapy-Splash Integration
- Align timeouts: Set
SPLASH_REQUEST_TIMEOUT = 30in your Scrapy settings to avoid orphaned requests lingering in memory. - Add a small delay: Set
DOWNLOAD_DELAY = 1to give Splash time to process and clean up between requests.
2. How to Add a 1-Minute Retry Delay for Docker Restarts
If you still need to rely on container restarts, you can customize Scrapy’s retry behavior to wait long enough for Docker to reboot. Here are two approaches:
Quick Settings Adjustment
Add these to your settings.py for a fixed 60-second delay between retries (disable exponential backoff for consistent waits):
RETRY_DELAY = 60 # 1-minute delay between retries RETRY_BACKOFF = 1 # Disable exponential backoff to keep delays fixed RETRY_TIMES = 3 # Allow enough retries to cover Docker's restart time
Custom Retry Middleware (Splash-Only Delays)
If you only want to delay retries for Splash requests (keep default behavior for others), create a custom middleware:
- Create
middlewares.pyin your Scrapy project with this code:
from scrapy.downloadermiddlewares.retry import RetryMiddleware from scrapy.utils.log import logger class DelayedSplashRetryMiddleware(RetryMiddleware): def __init__(self, settings): super().__init__(settings) self.splash_retry_delay = settings.getint('SPLASH_RETRY_DELAY', 60) def process_exception(self, request, exception, spider): # Check if this is a Splash request is_splash_request = 'splash' in request.meta.get('download_slot', '') or 'http://127.0.0.1:8050' in request.url if is_splash_request: retry_times = request.meta.get('retry_times', 0) + 1 if retry_times <= self.max_retry_times: logger.debug(f"Waiting {self.splash_retry_delay}s to retry Splash request (failed {retry_times}x)") # Schedule retry with delay return request.replace( meta=dict(request.meta, retry_times=retry_times), dont_filter=True, priority=request.priority - 1 ) logger.debug(f"Gave up on Splash request after {retry_times} retries") return None # Use default retry logic for non-Splash requests return super().process_exception(request, exception, spider)
- Update
settings.pyto replace the default retry middleware:
DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.retry.RetryMiddleware': None, 'your_project_name.middlewares.DelayedSplashRetryMiddleware': 550, } SPLASH_RETRY_DELAY = 60 RETRY_TIMES = 3
This ensures only Splash requests wait 1 minute before retrying, giving your Docker container time to restart fully.
内容的提问来源于stack exchange,提问作者Milano

