如何抓取op.gg中点击按钮后才显示的英雄联盟游戏统计数据?
Hey there! It sounds like you're stuck grabbing that extended game data that only shows up after clicking a button on OP.GG. Let's break down how to solve this—here are a couple of reliable approaches tailored to your setup:
1. Capture the AJAX Request Behind the Button Click
Most sites like OP.GG load additional content via AJAX (XHR/Fetch requests) when you interact with buttons, instead of reloading the whole page. Here's how to find and mimic that request:
- Open your browser's DevTools (F12), head to the Network tab, and filter by "XHR" or "Fetch".
- Click the button that expands the game data, then watch for a new request that pops up. This request will almost certainly return JSON containing the extended data you need.
- Check the request's URL, headers, and parameters—you’ll notice it uses values like
data-game-idanddata-summoner-idfrom the originalGameItemdiv (you already have access to these from your initial scrape!).
Example Scrapy Code Snippet
First, extract the required IDs from the initial page, then fire off the AJAX request:
def parse(self, response): # Loop through each game item in the default view for game_item in response.css('div.GameItem'): game_id = game_item.attrib.get('data-game-id') summoner_id = game_item.attrib.get('data-summoner-id') # Replace with the actual AJAX endpoint you found in DevTools ajax_url = f"https://op.gg/api/summoners/{summoner_id}/games/{game_id}/details" # Pass the base game data to the callback for combining later yield scrapy.Request( ajax_url, callback=self.parse_extended_data, meta={'base_game_data': { 'game_id': game_id, 'summoner_id': summoner_id, 'result': game_item.attrib.get('data-game-result') }} ) def parse_extended_data(self, response): # Parse the JSON response (adjust fields based on the actual API structure) extended_data = response.json() base_data = response.meta['base_game_data'] # Combine default and extended data into your final item yield { **base_data, 'kda': extended_data.get('kda'), 'full_item_list': extended_data.get('items'), 'wards_placed': extended_data.get('wards_placed'), # Add any other extended fields you need }
2. Simulate Browser Interaction (If AJAX Isn't an Option)
If the site renders the extended content using client-side JS without an explicit AJAX request, you’ll need to simulate a real browser click. Tools like Scrapy-Playwright work great for this, and it’s easy to integrate with Docker too.
Using Scrapy-Playwright
First, set up Scrapy-Playwright in your project (follow the official setup docs to install dependencies and configure settings). Then write a spider that clicks the button and waits for the extended content:
from scrapy_playwright.page import PageCoroutine def start_requests(self): yield scrapy.Request( url='YOUR_TARGET_SUMMONER_PROFILE_URL', meta={ 'playwright': True, 'playwright_page_coroutines': [ # Replace with the actual selector for the expand button PageCoroutine('click', selector='div.GameItem .expand-button'), # Wait for the extended class to appear before scraping PageCoroutine('wait_for_selector', selector='div.GameItem.extended'), ], } ) def parse(self, response): # Now you can scrape the extended content directly from the rendered page for extended_game in response.css('div.GameItem.extended'): yield { 'game_id': extended_game.attrib.get('data-game-id'), 'summoner_id': extended_game.attrib.get('data-summoner-id'), 'result': extended_game.attrib.get('data-game-result'), # Extract fields from the newly loaded extended section 'damage_dealt': extended_game.css('div.damage-stat::text').get(), 'vision_score': extended_game.css('div.vision-score::text').get(), }
Quick Tips to Avoid Blocking
- Always match the
User-Agentheader of a real browser to avoid being flagged. - Check OP.GG's terms of service to ensure your scraping activities are allowed.
- If using Docker, make sure your container includes all necessary dependencies for Playwright (like browser binaries).
内容的提问来源于stack exchange,提问作者DY Sung

