使用Scrapy或Selenium抓取页面时捕获后续HTTP请求
Absolutely, you can capture all those follow-up requests—whether they’re for images, CSS, scripts, or JS-triggered dynamic loads—using Scrapy + Splash or Selenium. Scapy would be overkill here, so let’s stick to simpler, more targeted tools that fit your web scraping workflow.
Option 1: Scrapy + Splash (Best for Scrapy Native Workflows)
Splash is a headless browser that integrates seamlessly with Scrapy, and it can capture every HTTP request made during page rendering (including those triggered by JavaScript). The key here is using Splash’s HAR (HTTP Archive) feature, which logs all network activity.
Setup & Basic Implementation:
- First, make sure you have Splash running (either via Docker or a local installation).
- Configure Scrapy to use Splash in your
settings.py:SPLASH_URL = 'http://localhost:8050' DOWNLOADER_MIDDLEWARES = { 'scrapy_splash.SplashCookiesMiddleware': 723, 'scrapy_splash.SplashMiddleware': 725, 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810, } SPIDER_MIDDLEWARES = { 'scrapy_splash.SplashDeduplicateArgsMiddleware': 100, } DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter' - In your spider, use
SplashRequestinstead ofRequest, and ask Splash to return the HAR data:import scrapy from scrapy_splash import SplashRequest class MySpider(scrapy.Spider): name = 'followup_crawler' start_urls = ['https://example.com'] def start_requests(self): for url in self.start_urls: yield SplashRequest( url, callback=self.parse, args={ 'wait': 3, # Wait 3s for JS to load dynamic content 'har': 1, # Enable HAR logging } ) def parse(self, response): # Extract all requests/responses from the HAR har = response.data['har'] for entry in har['log']['entries']: request_url = entry['request']['url'] response_status = entry['response']['status'] # Get response content (note: some content might be base64 encoded) response_content = entry['response']['content'].get('text', '') self.logger.info(f"Captured request: {request_url} | Status: {response_status}") # Do something with request_url and response_content here
This will capture every single HTTP request made after the initial GET, including embedded resources and JS-triggered calls.
Option 2: Selenium (Best for Highly Dynamic Pages)
If you prefer a more straightforward approach without setting up Splash, Selenium can capture network requests using the Chrome DevTools Protocol (CDP). You don’t need extra libraries (though selenium-wire can simplify it, but let’s stick to native Selenium first).
Basic Implementation with CDP:
from selenium import webdriver from selenium.webdriver.chrome.options import Options def capture_all_requests(url): chrome_options = Options() chrome_options.add_argument('--headless=new') # Run headless if needed driver = webdriver.Chrome(options=chrome_options) # Enable Network logging via CDP driver.execute_cdp_cmd('Network.enable', {}) # Store all requests and responses all_requests = [] # Listen for response received events def handle_response(event): request_url = event['request']['url'] response_status = event['response']['status'] # To get response body, you might need to use Network.getResponseBody # Note: This requires the request ID try: body = driver.execute_cdp_cmd('Network.getResponseBody', {'requestId': event['requestId']}) response_content = body.get('body', '') except: response_content = 'Unable to retrieve content' all_requests.append({ 'url': request_url, 'status': response_status, 'content': response_content }) # Add the listener driver.add_event_listener('Network.responseReceived', handle_response) # Load the page driver.get(url) # Wait for JS to finish loading (adjust time as needed) driver.implicitly_wait(5) # Print or process captured requests for req in all_requests: print(f"URL: {req['url']} | Status: {req['status']}") driver.quit() return all_requests # Usage capture_all_requests('https://example.com')
If you want even simpler code, selenium-wire is a wrapper that lets you access network requests directly without dealing with CDP events—just install it and use driver.requests to get all captured requests.
Which Should You Pick?
- Scrapy + Splash: Ideal if you’re already using Scrapy and want to integrate this into your existing pipeline. It’s faster than Selenium and fits Scrapy’s async model.
- Selenium: Better for one-off scripts or pages with extremely complex JavaScript (like single-page apps). It’s easier to debug visually if needed.
Both options avoid overengineering with Scapy, which is designed for low-level network packet manipulation—not web scraping follow-up requests.
内容的提问来源于stack exchange,提问作者Vasilis

