如何结合Scrapy与Selenium绕过年龄验证爬取Steam生化奇兵游戏数据
Hey there! Let's walk through how to combine Scrapy and Selenium exactly for your use case—handling those Steam age verification popups with Selenium only when needed, while letting Scrapy handle the rest of the scraping efficiently.
核心思路
We'll keep Scrapy as the workhorse for most requests and data parsing, and only call Selenium when we hit a page with an age verification popup. Once Selenium bypasses the check, we'll pass the processed page source back to Scrapy to extract the game details. This way you get the best of both worlds: Scrapy's speed and Selenium's ability to handle dynamic interactions.
Step-by-Step Implementation
1. Build a Custom Selenium Download Middleware
We'll create a Scrapy download middleware to detect age gates, use Selenium to get past them, and return the clean page to Scrapy.
Add this code to your project's middlewares.py file:
from scrapy import signals from scrapy.http import HtmlResponse from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time class SteamAgeVerificationMiddleware: def __init__(self): # Initialize headless Chrome (you can swap for Firefox if preferred) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') self.driver = webdriver.Chrome(options=chrome_options) def process_request(self, request, spider): # Only intercept Steam game app pages if "store.steampowered.com/app/" in request.url: self.driver.get(request.url) try: # Wait for age gate to appear (max 5 seconds) WebDriverWait(self.driver, 5).until( EC.presence_of_element_located((By.ID, "agegate_box")) ) # Bypass age check (set a birth year that's over 18) self.driver.find_element(By.ID, "ageYear").send_keys("1990") self.driver.find_element(By.ID, "view_product_page_btn").click() # Give the page a second to load after verification time.sleep(2) except: # No age gate found? Skip straight to returning the page pass # Pass the processed page source back to Scrapy page_source = self.driver.page_source return HtmlResponse( url=self.driver.current_url, body=page_source, encoding='utf-8', request=request ) # For non-game pages, let Scrapy handle the request normally return None def spider_closed(self, spider): # Clean up: close the browser when the spider finishes self.driver.quit()
2. Enable the Middleware
Update your project's settings.py to activate the middleware. Add this to the DOWNLOADER_MIDDLEWARES section:
DOWNLOADER_MIDDLEWARES = { 'your_project_name.middlewares.SteamAgeVerificationMiddleware': 543, # Keep other default middleware entries here—just make sure our middleware has a priority that runs before Scrapy's default downloader }
3. Write Your Scrapy Spider
Your spider logic can stay almost identical to what you already have—you just need to focus on parsing the game data, since the middleware handles the age gate. Here's an example:
import scrapy class BioshockSpider(scrapy.Spider): name = "bioshock" start_urls = [ "https://store.steampowered.com/search/?term=bioshock", ] def parse(self, response): # Grab all game links from the search results for game_card in response.css('a.search_result_row'): game_url = game_card.attrib['href'] yield scrapy.Request(game_url, callback=self.parse_game_details) def parse_game_details(self, response): # Extract game name, price, and URL game_name = response.css('div.apphub_AppName::text').get().strip() # Handle both regular and discounted prices game_price = response.css('div.game_purchase_price::text').get() if not game_price: game_price = response.css('div.discount_final_price::text').get() game_price = game_price.strip() if game_price else "Price unavailable" # Yield the data (Scrapy will handle saving to CSV if you set up feeds) yield { 'name': game_name, 'price': game_price, 'url': response.url }
4. Optimize CSV Saving (Optional but Recommended)
Instead of writing CSV manually, use Scrapy's built-in Feed Exports. Add this to settings.py:
FEEDS = { 'bioshock_games.csv': { 'format': 'csv', 'fields': ['name', 'price', 'url'], 'overwrite': True, 'encoding': 'utf-8', }, }
Quick Tips to Avoid Headaches
- Performance: Selenium is slower than Scrapy, so we only use it for age-gated pages—don't route all requests through it!
- Anti-Crawling: Steam has anti-bot measures, so add a delay in
settings.pylikeDOWNLOAD_DELAY = 2to avoid getting blocked. - Driver Compatibility: Make sure your ChromeDriver (or other browser driver) version matches your installed browser.
Test It Out
Run your spider with:
scrapy crawl bioshock
It should automatically bypass age verification for any Bioshock games that require it, and save all the game names, prices, and links to your CSV file.
内容的提问来源于stack exchange,提问作者wuggs

