You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取op.gg中点击按钮后才显示的英雄联盟游戏统计数据?

How to Scrape Click-to-Load Extended Game Data from OP.GG with Scrapy

Hey there! It sounds like you're stuck grabbing that extended game data that only shows up after clicking a button on OP.GG. Let's break down how to solve this—here are a couple of reliable approaches tailored to your setup:

1. Capture the AJAX Request Behind the Button Click

Most sites like OP.GG load additional content via AJAX (XHR/Fetch requests) when you interact with buttons, instead of reloading the whole page. Here's how to find and mimic that request:

  • Open your browser's DevTools (F12), head to the Network tab, and filter by "XHR" or "Fetch".
  • Click the button that expands the game data, then watch for a new request that pops up. This request will almost certainly return JSON containing the extended data you need.
  • Check the request's URL, headers, and parameters—you’ll notice it uses values like data-game-id and data-summoner-id from the original GameItem div (you already have access to these from your initial scrape!).

Example Scrapy Code Snippet

First, extract the required IDs from the initial page, then fire off the AJAX request:

def parse(self, response):
    # Loop through each game item in the default view
    for game_item in response.css('div.GameItem'):
        game_id = game_item.attrib.get('data-game-id')
        summoner_id = game_item.attrib.get('data-summoner-id')
        
        # Replace with the actual AJAX endpoint you found in DevTools
        ajax_url = f"https://op.gg/api/summoners/{summoner_id}/games/{game_id}/details"
        
        # Pass the base game data to the callback for combining later
        yield scrapy.Request(
            ajax_url,
            callback=self.parse_extended_data,
            meta={'base_game_data': {
                'game_id': game_id,
                'summoner_id': summoner_id,
                'result': game_item.attrib.get('data-game-result')
            }}
        )

def parse_extended_data(self, response):
    # Parse the JSON response (adjust fields based on the actual API structure)
    extended_data = response.json()
    base_data = response.meta['base_game_data']
    
    # Combine default and extended data into your final item
    yield {
        **base_data,
        'kda': extended_data.get('kda'),
        'full_item_list': extended_data.get('items'),
        'wards_placed': extended_data.get('wards_placed'),
        # Add any other extended fields you need
    }

2. Simulate Browser Interaction (If AJAX Isn't an Option)

If the site renders the extended content using client-side JS without an explicit AJAX request, you’ll need to simulate a real browser click. Tools like Scrapy-Playwright work great for this, and it’s easy to integrate with Docker too.

Using Scrapy-Playwright

First, set up Scrapy-Playwright in your project (follow the official setup docs to install dependencies and configure settings). Then write a spider that clicks the button and waits for the extended content:

from scrapy_playwright.page import PageCoroutine

def start_requests(self):
    yield scrapy.Request(
        url='YOUR_TARGET_SUMMONER_PROFILE_URL',
        meta={
            'playwright': True,
            'playwright_page_coroutines': [
                # Replace with the actual selector for the expand button
                PageCoroutine('click', selector='div.GameItem .expand-button'),
                # Wait for the extended class to appear before scraping
                PageCoroutine('wait_for_selector', selector='div.GameItem.extended'),
            ],
        }
    )

def parse(self, response):
    # Now you can scrape the extended content directly from the rendered page
    for extended_game in response.css('div.GameItem.extended'):
        yield {
            'game_id': extended_game.attrib.get('data-game-id'),
            'summoner_id': extended_game.attrib.get('data-summoner-id'),
            'result': extended_game.attrib.get('data-game-result'),
            # Extract fields from the newly loaded extended section
            'damage_dealt': extended_game.css('div.damage-stat::text').get(),
            'vision_score': extended_game.css('div.vision-score::text').get(),
        }

Quick Tips to Avoid Blocking

  • Always match the User-Agent header of a real browser to avoid being flagged.
  • Check OP.GG's terms of service to ensure your scraping activities are allowed.
  • If using Docker, make sure your container includes all necessary dependencies for Playwright (like browser binaries).

内容的提问来源于stack exchange,提问作者DY Sung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:21:48