You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言爬取英超官网动态统计数据的技术问询

Solution for Dynamic Premier League Stats Scraping in R

Hey there! I've dealt with this exact dynamic scraping scenario before—since the Premier League stats page loads data via JavaScript (no URL changes when switching seasons or pages), we can't rely solely on rvest for static HTML. Instead, we'll mimic the AJAX requests the page makes under the hood to pull data directly from their API.

Step 1: Inspect the Page's API Requests

First, open your browser's DevTools (F12 > Network tab). When you switch seasons or click "Next" for more players, you'll see an XHR request pop up. This request is fetching the dynamic data, and it will have parameters that map to the filters you set:

  • Your data-metric corresponds to the statType parameter (e.g., "goals" for your use case)
  • data-page maps to the page parameter (starts at 1, increments by 1 each page)
  • You'll also spot a compSeasons parameter for the season ID (e.g., "274" = 2023/24 season)

Step 2: Code to Fetch Dynamic Data

We'll use httr to send requests to the API and jsonlite to parse the JSON response, alongside tidyverse for data wrangling. Here's a reusable function:

library(tidyverse)
library(httr)
library(jsonlite)

get_pl_player_stats <- function(season_id, stat_type = "goals", page = 1, page_size = 20) {
  # The API endpoint the website uses for stats
  api_url <- "https://footballapi.pulselive.com/football/stats/top"
  
  # Request parameters matching the page's filters
  request_params <- list(
    compSeasons = season_id,
    statType = stat_type,
    page = page,
    pageSize = page_size,
    competitionId = "PL",
    altIds = "true",
    detail = "false"
  )
  
  # Send the GET request with headers to avoid being blocked
  response <- GET(
    url = api_url,
    query = request_params,
    add_headers(
      "Origin" = "https://www.premierleague.com",
      "Referer" = "https://www.premierleague.com/stats/top/players/goals"
    )
  )
  
  # Parse the JSON response into a usable data frame
  raw_data <- fromJSON(content(response, "text"))
  
  # Extract and clean the relevant player stats
  cleaned_stats <- raw_data$stats$content %>%
    as_tibble() %>%
    select(
      player_name = player.name,
      player_id = player.id,
      stat_value = value,
      team_name = team.name,
      season = compSeason.label
    )
  
  return(cleaned_stats)
}

Step 3: Use the Function

  • Single page/season: Fetch the first page of 2023/24 season goals (season ID "274"):
    single_page_data <- get_pl_player_stats(season_id = "274", page = 1)
    
  • Multiple pages: Fetch the first 3 pages of data using purrr::map_dfr:
    all_pages <- 1:3
    multi_page_data <- map_dfr(all_pages, ~get_pl_player_stats(season_id = "274", page = .x))
    
  • Different seasons: Find the season ID via DevTools (switch seasons and check the XHR parameters) and plug it in. For example, "266" corresponds to the 2022/23 season.

Why This Works

Instead of scraping the static HTML table, we're pulling data directly from the same API the website uses. This avoids headaches with JavaScript-rendered content and lets us easily control filters like season, page, and stat type.

内容的提问来源于stack exchange,提问作者CFlaherty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:17:40