基于R语言爬取英超官网动态统计数据的技术问询
Hey there! I've dealt with this exact dynamic scraping scenario before—since the Premier League stats page loads data via JavaScript (no URL changes when switching seasons or pages), we can't rely solely on rvest for static HTML. Instead, we'll mimic the AJAX requests the page makes under the hood to pull data directly from their API.
Step 1: Inspect the Page's API Requests
First, open your browser's DevTools (F12 > Network tab). When you switch seasons or click "Next" for more players, you'll see an XHR request pop up. This request is fetching the dynamic data, and it will have parameters that map to the filters you set:
- Your
data-metriccorresponds to thestatTypeparameter (e.g., "goals" for your use case) data-pagemaps to thepageparameter (starts at 1, increments by 1 each page)- You'll also spot a
compSeasonsparameter for the season ID (e.g., "274" = 2023/24 season)
Step 2: Code to Fetch Dynamic Data
We'll use httr to send requests to the API and jsonlite to parse the JSON response, alongside tidyverse for data wrangling. Here's a reusable function:
library(tidyverse) library(httr) library(jsonlite) get_pl_player_stats <- function(season_id, stat_type = "goals", page = 1, page_size = 20) { # The API endpoint the website uses for stats api_url <- "https://footballapi.pulselive.com/football/stats/top" # Request parameters matching the page's filters request_params <- list( compSeasons = season_id, statType = stat_type, page = page, pageSize = page_size, competitionId = "PL", altIds = "true", detail = "false" ) # Send the GET request with headers to avoid being blocked response <- GET( url = api_url, query = request_params, add_headers( "Origin" = "https://www.premierleague.com", "Referer" = "https://www.premierleague.com/stats/top/players/goals" ) ) # Parse the JSON response into a usable data frame raw_data <- fromJSON(content(response, "text")) # Extract and clean the relevant player stats cleaned_stats <- raw_data$stats$content %>% as_tibble() %>% select( player_name = player.name, player_id = player.id, stat_value = value, team_name = team.name, season = compSeason.label ) return(cleaned_stats) }
Step 3: Use the Function
- Single page/season: Fetch the first page of 2023/24 season goals (season ID "274"):
single_page_data <- get_pl_player_stats(season_id = "274", page = 1) - Multiple pages: Fetch the first 3 pages of data using
purrr::map_dfr:all_pages <- 1:3 multi_page_data <- map_dfr(all_pages, ~get_pl_player_stats(season_id = "274", page = .x)) - Different seasons: Find the season ID via DevTools (switch seasons and check the XHR parameters) and plug it in. For example, "266" corresponds to the 2022/23 season.
Why This Works
Instead of scraping the static HTML table, we're pulling data directly from the same API the website uses. This avoids headaches with JavaScript-rendered content and lets us easily control filters like season, page, and stat type.
内容的提问来源于stack exchange,提问作者CFlaherty

