Pro Football Reference多页爬虫报错求助:无法加载外部实体
解决建议
修复URL协议问题:当前使用的
http协议已被目标站点弃用,该网站强制使用HTTPS。将调用语句中的URL前缀改为https://www.pro-football-reference.com/years/,避免重定向导致的加载失败。添加请求头绕过反爬拦截:网站会拦截无标识的爬虫请求,需模拟浏览器发送请求。可以结合
httr包先获取页面内容,再用XML包解析:library(XML) library(httr) scrapeData = function(urlprefix, urlend, startyr, endyr) { master = data.frame() # 设置浏览器请求头 headers = add_headers( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" ) for (i in startyr:endyr) { cat('Loading Year', i, '\n') URL = paste(urlprefix, as.character(i), urlend, sep = "") # 先发送请求获取页面内容 response = GET(URL, headers) # 解析HTML内容 table = readHTMLTable(content(response, "text"), stringsAsFactors = F)[[1]] table$Year = i master = rbind(table, master) } return(master) } # 调用时使用HTTPS前缀 drafts = scrapeData('https://www.pro-football-reference.com/years/', '/draft.htm', 2010, 2010)改用更稳定的rvest包:
XML包的readHTMLTable对现代网页支持有限,推荐使用rvest包(tidyverse生态工具),处理表格更可靠:library(rvest) library(dplyr) scrapeData = function(urlprefix, urlend, startyr, endyr) { master = tibble() headers = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" for (i in startyr:endyr) { cat('Loading Year', i, '\n') URL = paste(urlprefix, as.character(i), urlend, sep = "") # 读取页面并提取第一个表格 table = read_html(URL, user_agent = headers) %>% html_table(header = TRUE, stringsAsFactors = FALSE) %>% .[[1]] %>% mutate(Year = i) master = bind_rows(master, table) } return(master) } drafts = scrapeData('https://www.pro-football-reference.com/years/', '/draft.htm', 2010, 2010)额外提示:避免短时间内频繁请求,可在循环中添加
Sys.sleep(1)设置1秒延迟,防止被网站封禁IP。
内容的提问来源于stack exchange,提问作者Daniel
相关产品推荐
相关产品推荐

