You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R爬取Understat英超2020赛季球队数据遇到的问题:数据行数异常及表格获取失败

Fixing Understat EPL 2020 Data Scraping Issues

Let's break down what's going on here and fix your problems step by step:

Why You're Getting 600 Rows Instead of 20

The x$history field in the parsed JSON contains game-by-game match data for each team, not the season-wide summary stats you're expecting. For the 2020/21 EPL season, each team plays 38 matches, so combining all those single-game entries for 20 teams gives you the 600 rows you're seeing. To get one row per team, you need to use the x$total field instead—it stores aggregated season totals for every team.

Revised Code to Pull Season Summary Data

Here's the updated code that correctly extracts the 20-row season summary, with more reliable JSON parsing:

library(rvest)
library(stringr)
library(jsonlite)
library(dplyr)

# Fetch the page
xg_page <- "https://understat.com/league/EPL/2020"
xg_scraped_page <- read_html(xg_page)

# Extract the script containing teamsData
xg_script <- xg_scraped_page %>% 
  html_nodes("script") %>% 
  str_subset("teamsData") %>% 
  first() # Ensure we only grab the relevant script tag

# Extract and unescape the JSON string
xg_json <- str_extract(xg_script, "(?<=JSON.parse\\(').*(?='\\);)") %>% 
  stringi::stri_unescape_unicode()

# Parse the JSON into a list
teams_data <- fromJSON(xg_json, simplifyDataFrame = TRUE, flatten = TRUE)

# Extract season-wide totals (one row per team)
teams_summary <- lapply(teams_data, function(team) {
  team_total <- team$total
  team_total$team_id <- team$id
  team_total$team_name <- team$title
  team_total
}) %>% 
  bind_rows() # Use dplyr's bind_rows for cleaner data frame binding

# Build your desired xG/xGA summary table
teams_xg_xga <- teams_summary %>% 
  select(team_name, xG, xGA) %>% 
  rename(teams = team_name)

# Verify row count (should be 20)
nrow(teams_xg_xga)

Key Changes:

  • Replaced x$history with x$total: This pulls aggregated season stats instead of single-game data.
  • Improved JSON extraction: Used str_extract with a precise regex to target the exact JSON string inside the script tag, avoiding potential mismatches from the original sub call.
  • Used bind_rows instead of do.call(rbind, ...): A more robust and readable way to combine data frames in dplyr.

Accessing the "League Champ" (Standings) Table

I assume "league chemp" refers to the main league standings table on the page. The teams_summary data frame we created already contains all the data needed for this table (points, wins, draws, goals scored, etc.). You can build a full, sorted standings table with this code:

league_standings <- teams_summary %>% 
  select(team_name, pts, wins, draws, loses, scored, missed, xG, xGA) %>% 
  arrange(desc(pts)) # Sort by points to match the page's standings order

This will give you a table identical to the one displayed on the Understat page, with one row per team.

内容的提问来源于stack exchange,提问作者AW27

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 13:37:31