使用R爬取Understat英超2020赛季球队数据遇到的问题:数据行数异常及表格获取失败
Let's break down what's going on here and fix your problems step by step:
Why You're Getting 600 Rows Instead of 20
The x$history field in the parsed JSON contains game-by-game match data for each team, not the season-wide summary stats you're expecting. For the 2020/21 EPL season, each team plays 38 matches, so combining all those single-game entries for 20 teams gives you the 600 rows you're seeing. To get one row per team, you need to use the x$total field instead—it stores aggregated season totals for every team.
Revised Code to Pull Season Summary Data
Here's the updated code that correctly extracts the 20-row season summary, with more reliable JSON parsing:
library(rvest) library(stringr) library(jsonlite) library(dplyr) # Fetch the page xg_page <- "https://understat.com/league/EPL/2020" xg_scraped_page <- read_html(xg_page) # Extract the script containing teamsData xg_script <- xg_scraped_page %>% html_nodes("script") %>% str_subset("teamsData") %>% first() # Ensure we only grab the relevant script tag # Extract and unescape the JSON string xg_json <- str_extract(xg_script, "(?<=JSON.parse\\(').*(?='\\);)") %>% stringi::stri_unescape_unicode() # Parse the JSON into a list teams_data <- fromJSON(xg_json, simplifyDataFrame = TRUE, flatten = TRUE) # Extract season-wide totals (one row per team) teams_summary <- lapply(teams_data, function(team) { team_total <- team$total team_total$team_id <- team$id team_total$team_name <- team$title team_total }) %>% bind_rows() # Use dplyr's bind_rows for cleaner data frame binding # Build your desired xG/xGA summary table teams_xg_xga <- teams_summary %>% select(team_name, xG, xGA) %>% rename(teams = team_name) # Verify row count (should be 20) nrow(teams_xg_xga)
Key Changes:
- Replaced
x$historywithx$total: This pulls aggregated season stats instead of single-game data. - Improved JSON extraction: Used
str_extractwith a precise regex to target the exact JSON string inside the script tag, avoiding potential mismatches from the originalsubcall. - Used
bind_rowsinstead ofdo.call(rbind, ...): A more robust and readable way to combine data frames in dplyr.
Accessing the "League Champ" (Standings) Table
I assume "league chemp" refers to the main league standings table on the page. The teams_summary data frame we created already contains all the data needed for this table (points, wins, draws, goals scored, etc.). You can build a full, sorted standings table with this code:
league_standings <- teams_summary %>% select(team_name, pts, wins, draws, loses, scored, missed, xG, xGA) %>% arrange(desc(pts)) # Sort by points to match the page's standings order
This will give you a table identical to the one displayed on the Understat page, with one row per team.
内容的提问来源于stack exchange,提问作者AW27

