R语言批量获取Baseball Prospectus数据遇错及文件覆盖问题求助
嗨,我来帮你一步步解决这两个R语言的问题:
问题1:HTML抓取时的
list_to_dataframe报错 这个错误的核心原因是:fetch_adjusted函数在处理某些日期时,返回的结果不是标准的数据框(比如页面解析失败返回了NULL,或者某天的表格结构和其他日期不一致)。而ldply要求所有函数返回的结果要么都是原子向量,要么都是数据框,类型不统一就会触发这个报错。
我给你调整了代码,添加了错误捕获机制,确保哪怕单个日期处理失败,整个批量任务也能继续,且始终返回结构一致的数据框:
library(plyr) library(XML) fetch_adjusted <- function(day) { # 把日期格式化为两位数字(比如1变成01),避免URL格式不规范 day_str <- sprintf("%02d", day) fname <- paste0("standings201909", day_str, ".html") url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=2019-09-", day_str) # 捕获下载过程中的错误 download_success <- tryCatch({ download.file(url = url, destfile = fname) TRUE }, error = function(e) { message(paste("下载", day_str, "日数据失败:", e$message)) FALSE }) if (!download_success) { # 返回和正常结果结构一致的空数据框 return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0))) } # 捕获HTML解析错误 doc0 <- tryCatch(htmlParse(file = fname, encoding = "UTF-8"), error = function(e) { message(paste("解析", day_str, "日HTML失败:", e$message)) NULL }) if (is.null(doc0)) { return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0))) } doc1 <- xmlRoot(doc0) doc2 <- getNodeSet(doc1, "//table[@id='content']") if (length(doc2) == 0) { message(paste(day_str, "日未找到目标表格")) return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0))) } # 捕获表格读取错误 standings <- tryCatch(readHTMLTable(doc2[[1]], header = TRUE, skip.rows = 1, stringsAsFactors = FALSE), error = function(e) { message(paste("读取", day_str, "日表格失败:", e$message)) NULL }) if (is.null(standings) || length(standings) == 0) { return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0))) } standings <- standings[[1]] standings$day <- day_str standings } # 先获取一个正常的模板数据框,用来定义空数据框的列结构 standings_template <- fetch_adjusted(1) # 批量处理9月1-29日的数据 Sept <- ldply(1:29, fetch_adjusted, .progress = "text")
代码说明:
- 用
sprintf格式化日期为两位数字,保证URL格式统一 - 多层
tryCatch捕获下载、解析、读取表格各环节的错误,避免单个日期失败导致整个任务中断 - 定义
standings_template确保空数据框和正常结果列数一致,让ldply能顺利合并
问题2:XLS文件被反复覆盖
这个问题很直观:你每次下载都用同一个文件名test.xls,新文件会直接覆盖旧文件,最后自然只剩最后一次的结果。另外你提到的“空表格”,大概率是最后一天的XLS本身为空,或者下载时没正确用二进制模式(wb是XLS文件的正确下载模式)。
给你两个解决方案,按需选择:
方案1:保存为多个独立的XLS文件(每个日期一个)
修改文件名,让每个文件包含对应的日期,这样就不会互相覆盖:
dates <- seq(as.Date("2019-09-01"), as.Date("2019-09-30"), by = 1) fetch_adjusted <- function(date) { date_str <- as.character(date) url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=", date_str, "&otype=xls") # 生成带日期的唯一文件名 destfile <- paste0("standings_", date_str, ".xls") # 捕获下载错误,用wb模式下载二进制文件 tryCatch({ download.file(url, destfile, mode = "wb") message(paste("成功下载:", destfile)) }, error = function(e) { message(paste("下载", date_str, "失败:", e$message)) }) } # 批量下载所有日期的文件 lapply(dates, fetch_adjusted)
方案2:直接读取并合并成一个数据框(不保留单个文件)
如果不需要保存原始XLS文件,而是直接得到合并后的数据集,可以用临时文件来处理:
library(plyr) library(readxl) # 需要先安装:install.packages("readxl") dates <- seq(as.Date("2019-09-01"), as.Date("2019-09-30"), by = 1) fetch_adjusted <- function(date) { date_str <- as.character(date) url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=", date_str, "&otype=xls") # 创建临时文件,下载后读取就删除 temp_file <- tempfile(fileext = ".xls") # 捕获下载错误 download_success <- tryCatch({ download.file(url, temp_file, mode = "wb") TRUE }, error = function(e) { message(paste("下载", date_str, "失败:", e$message)) FALSE }) if (!download_success) { return(data.frame()) } # 读取XLS文件,可根据实际表格结构调整skip、col_names等参数 standings <- tryCatch(read_excel(temp_file), error = function(e) { message(paste("读取", date_str, "数据失败:", e$message)) data.frame() }) # 删除临时文件 file.remove(temp_file) # 添加日期标识列 standings$date <- date_str standings } # 批量处理并合并成一个大的数据框 combined_standings <- ldply(dates, fetch_adjusted, .progress = "text")
方案说明:
- 方案1适合需要保留原始文件归档的场景,每个文件有唯一的日期命名
- 方案2适合直接做数据分析的场景,用临时文件避免磁盘冗余,一步得到合并后的数据集
内容的提问来源于stack exchange,提问作者Abb
相关产品推荐
相关产品推荐

