R语言抓取Baseball Reference投球数据表格失败求助
Solution for Scraping Baseball Reference Pitching Table
The issue with your pitching table scrape is likely due to targeting the table element directly instead of its wrapping container (matching how you successfully scraped the batting table). Here's the fixed code:
url2017 <- "https://www.baseball-reference.com/leagues/majors/2017.shtml" webpage2017 <- read_html(url2017) # Scrape batting data (unchanged) hitting_data2017 <- webpage2017 %>% html_nodes('#div_teams_standard_batting') %>% html_table() hitterdata2017 <- hitting_data2017[[1]] # Scrape pitching data (fixed selector to use the wrapping div) pitching_data2017 <- webpage2017 %>% html_nodes('#div_teams_standard_pitching') %>% html_table() pitchingdata2017 <- pitching_data2017[[1]] # Optional: Clean up repeated header rows in both tables hitterdata2017 <- hitterdata2017 %>% filter(Team != "Team") pitchingdata2017 <- pitchingdata2017 %>% filter(Team != "Team")
Key Fix Explanation:
- Baseball Reference wraps each stats table in a
<div>with an ID formatted asdiv_<table_id>. For pitching, this isdiv_teams_standard_pitching. - Targeting this div ensures
html_table()correctly parses the embedded table, avoiding potential issues with the table's internal header structure that might block direct table selection.
If you still encounter issues, ensure your rvest package is up to date by running install.packages("rvest").
内容的提问来源于stack exchange,提问作者pcm1113
相关产品推荐
相关产品推荐

