You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy或Selenium抓取页面时捕获后续HTTP请求

Absolutely, you can capture all those follow-up requests—whether they’re for images, CSS, scripts, or JS-triggered dynamic loads—using Scrapy + Splash or Selenium. Scapy would be overkill here, so let’s stick to simpler, more targeted tools that fit your web scraping workflow.

Option 1: Scrapy + Splash (Best for Scrapy Native Workflows)

Splash is a headless browser that integrates seamlessly with Scrapy, and it can capture every HTTP request made during page rendering (including those triggered by JavaScript). The key here is using Splash’s HAR (HTTP Archive) feature, which logs all network activity.

Setup & Basic Implementation:

  1. First, make sure you have Splash running (either via Docker or a local installation).
  2. Configure Scrapy to use Splash in your settings.py:
    SPLASH_URL = 'http://localhost:8050'
    DOWNLOADER_MIDDLEWARES = {
        'scrapy_splash.SplashCookiesMiddleware': 723,
        'scrapy_splash.SplashMiddleware': 725,
        'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
    }
    SPIDER_MIDDLEWARES = {
        'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
    }
    DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'
    
  3. In your spider, use SplashRequest instead of Request, and ask Splash to return the HAR data:
    import scrapy
    from scrapy_splash import SplashRequest
    
    class MySpider(scrapy.Spider):
        name = 'followup_crawler'
        start_urls = ['https://example.com']
    
        def start_requests(self):
            for url in self.start_urls:
                yield SplashRequest(
                    url,
                    callback=self.parse,
                    args={
                        'wait': 3,  # Wait 3s for JS to load dynamic content
                        'har': 1,   # Enable HAR logging
                    }
                )
    
        def parse(self, response):
            # Extract all requests/responses from the HAR
            har = response.data['har']
            for entry in har['log']['entries']:
                request_url = entry['request']['url']
                response_status = entry['response']['status']
                # Get response content (note: some content might be base64 encoded)
                response_content = entry['response']['content'].get('text', '')
                
                self.logger.info(f"Captured request: {request_url} | Status: {response_status}")
                # Do something with request_url and response_content here
    

This will capture every single HTTP request made after the initial GET, including embedded resources and JS-triggered calls.

Option 2: Selenium (Best for Highly Dynamic Pages)

If you prefer a more straightforward approach without setting up Splash, Selenium can capture network requests using the Chrome DevTools Protocol (CDP). You don’t need extra libraries (though selenium-wire can simplify it, but let’s stick to native Selenium first).

Basic Implementation with CDP:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def capture_all_requests(url):
    chrome_options = Options()
    chrome_options.add_argument('--headless=new')  # Run headless if needed

    driver = webdriver.Chrome(options=chrome_options)
    # Enable Network logging via CDP
    driver.execute_cdp_cmd('Network.enable', {})
    
    # Store all requests and responses
    all_requests = []
    
    # Listen for response received events
    def handle_response(event):
        request_url = event['request']['url']
        response_status = event['response']['status']
        # To get response body, you might need to use Network.getResponseBody
        # Note: This requires the request ID
        try:
            body = driver.execute_cdp_cmd('Network.getResponseBody', {'requestId': event['requestId']})
            response_content = body.get('body', '')
        except:
            response_content = 'Unable to retrieve content'
        
        all_requests.append({
            'url': request_url,
            'status': response_status,
            'content': response_content
        })
    
    # Add the listener
    driver.add_event_listener('Network.responseReceived', handle_response)
    
    # Load the page
    driver.get(url)
    # Wait for JS to finish loading (adjust time as needed)
    driver.implicitly_wait(5)
    
    # Print or process captured requests
    for req in all_requests:
        print(f"URL: {req['url']} | Status: {req['status']}")
    
    driver.quit()
    return all_requests

# Usage
capture_all_requests('https://example.com')

If you want even simpler code, selenium-wire is a wrapper that lets you access network requests directly without dealing with CDP events—just install it and use driver.requests to get all captured requests.

Which Should You Pick?

  • Scrapy + Splash: Ideal if you’re already using Scrapy and want to integrate this into your existing pipeline. It’s faster than Selenium and fits Scrapy’s async model.
  • Selenium: Better for one-off scripts or pages with extremely complex JavaScript (like single-page apps). It’s easier to debug visually if needed.

Both options avoid overengineering with Scapy, which is designed for low-level network packet manipulation—not web scraping follow-up requests.

内容的提问来源于stack exchange,提问作者Vasilis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:14:29