如何在R语言中爬取网页表格?MLB数据爬取报错求助
Hey there! The error you're seeing (Error in .[[1]] : subscript out of bounds) happens because the table on that MLB stats page is dynamically rendered with JavaScript—the static HTML you get from read_html() doesn't actually contain the table element yet. Let's break down two solid solutions to fix this:
Most modern sports stats sites load data via backend APIs instead of embedding it directly in static HTML. For MLB, you can bypass the dynamic page entirely and pull data straight from their API, which is way more efficient.
Here's how to do it with httr and jsonlite:
library(jsonlite) library(httr) # Define the API endpoint and parameters matching your original request api_url <- "https://bdfed.stitch.mlbinfra.com/bdfed/stats/player" params <- list( sportId = "1", gameType = "R", season = "2018", seasonType = "ANY", leagueId = "103,104", # Corresponding to MLB's American and National leagues statType = "hitting", playerPool = "QUALIFIER", sortStat = "avg", sortOrder = "desc", perPage = "50", page = "1" ) # Send the request and parse the JSON response response <- GET(api_url, query = params) raw_data <- fromJSON(content(response, "text")) # Convert the stats data to a data frame hitting2018 <- as.data.frame(raw_data$stats)
This will give you the exact hitting stats table you want, without dealing with dynamic page rendering.
If you prefer to mimic a real user's browser behavior (good for trickier dynamic pages), you can use RSelenium to load the page fully, including JavaScript-rendered content.
Note: You'll need to install a browser driver first (like ChromeDriver for Chrome, or GeckoDriver for Firefox).
library(RSelenium) library(rvest) # Start a Chrome driver (adjust browser/port as needed) driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # Navigate to the MLB stats page target_url <- "http://mlb.mlb.com/stats/sortable.jsp#elem=%5Bobject+Object%5D&tab_level=child&click_text=Sortable+Player+hitting&game_type='R'&season=2018&season_type=ANY&league_code='MLB'§ionType=sp&statType=hitting&page=1&ts=1567176051240&playerType=QUALIFIER&sportCode='mlb'&split=&team_id=&active_sw=&position=&page_type=SortablePlayer&sortOrder='desc'&sortColumn=avg&results=&perPage=50&timeframe=&last_x_days=&extended=0" remDr$navigate(target_url) # Wait 3 seconds for the page to fully load (adjust if needed) Sys.sleep(3) # Grab the fully rendered page source and extract the table page_source <- remDr$getPageSource()[[1]] data <- read_html(page_source) hitting2018 <- data %>% html_nodes("table") %>% html_table(fill = TRUE) %>% .[[1]] # Clean up: close the browser and stop the driver remDr$close() driver$server$stop()
Quick Note:
The API method is always preferred—it's faster, more stable, and less likely to break if MLB updates their page design. The RSelenium method is a fallback for cases where you can't find the underlying API.
内容的提问来源于stack exchange,提问作者Dany

