You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy-splash内存占用过高问题及重试请求延迟配置咨询

Scrapy-Splash Memory Leak & Retry Delay Solutions

Let’s break this down step by step—first fixing the root memory issue (since 3 pages shouldn’t hog 10GB), then addressing the retry delay if you still need it as a fallback.

1. Better Solutions to Resolve the Memory Leak

Tune Splash Container Configuration

Splash is likely the source of the leak, even with low concurrent requests from Scrapy:

  • Match Splash’s render slots to Scrapy’s concurrency: By default, Splash uses more slots than you might expect. When starting the container, add --slots 3 to align with your CONCURRENT_REQUESTS=3 setting. This prevents Splash from overloading itself with unneeded render processes:
    docker run -p 8050:8050 --slots 3 scrapinghub/splash
    
  • Set hard memory limits for Docker: Cap the container’s memory to avoid crashing your entire system. Use --memory and --memory-swap to trigger a clean OOM kill, plus --restart=always to auto-restart the container without manual intervention:
    docker run -p 8050:8050 --slots 3 --memory=2g --memory-swap=2g --restart=always scrapinghub/splash
    
  • Limit render timeouts: Add --max-timeout 30 to prevent stuck renders from hanging around and consuming memory indefinitely.

Audit Custom Lua Scripts (If Used)

If you’re using a custom Lua script with Splash, make sure you’re cleaning up resources:

  • Always call splash:stop() after rendering a page to close the browser session and free memory.
  • Avoid storing large, unused objects in the script’s scope that won’t get garbage collected.

Optimize Scrapy-Splash Integration

  • Align timeouts: Set SPLASH_REQUEST_TIMEOUT = 30 in your Scrapy settings to avoid orphaned requests lingering in memory.
  • Add a small delay: Set DOWNLOAD_DELAY = 1 to give Splash time to process and clean up between requests.

2. How to Add a 1-Minute Retry Delay for Docker Restarts

If you still need to rely on container restarts, you can customize Scrapy’s retry behavior to wait long enough for Docker to reboot. Here are two approaches:

Quick Settings Adjustment

Add these to your settings.py for a fixed 60-second delay between retries (disable exponential backoff for consistent waits):

RETRY_DELAY = 60  # 1-minute delay between retries
RETRY_BACKOFF = 1  # Disable exponential backoff to keep delays fixed
RETRY_TIMES = 3  # Allow enough retries to cover Docker's restart time

Custom Retry Middleware (Splash-Only Delays)

If you only want to delay retries for Splash requests (keep default behavior for others), create a custom middleware:

  1. Create middlewares.py in your Scrapy project with this code:
from scrapy.downloadermiddlewares.retry import RetryMiddleware
from scrapy.utils.log import logger

class DelayedSplashRetryMiddleware(RetryMiddleware):
    def __init__(self, settings):
        super().__init__(settings)
        self.splash_retry_delay = settings.getint('SPLASH_RETRY_DELAY', 60)

    def process_exception(self, request, exception, spider):
        # Check if this is a Splash request
        is_splash_request = 'splash' in request.meta.get('download_slot', '') or 'http://127.0.0.1:8050' in request.url
        
        if is_splash_request:
            retry_times = request.meta.get('retry_times', 0) + 1
            if retry_times <= self.max_retry_times:
                logger.debug(f"Waiting {self.splash_retry_delay}s to retry Splash request (failed {retry_times}x)")
                # Schedule retry with delay
                return request.replace(
                    meta=dict(request.meta, retry_times=retry_times),
                    dont_filter=True,
                    priority=request.priority - 1
                )
            logger.debug(f"Gave up on Splash request after {retry_times} retries")
            return None
        
        # Use default retry logic for non-Splash requests
        return super().process_exception(request, exception, spider)
  1. Update settings.py to replace the default retry middleware:
DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.retry.RetryMiddleware': None,
    'your_project_name.middlewares.DelayedSplashRetryMiddleware': 550,
}
SPLASH_RETRY_DELAY = 60
RETRY_TIMES = 3

This ensures only Splash requests wait 1 minute before retrying, giving your Docker container time to restart fully.

内容的提问来源于stack exchange,提问作者Milano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:47:18