You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言批量获取Baseball Prospectus数据遇错及文件覆盖问题求助

嗨,我来帮你一步步解决这两个R语言的问题:

问题1:HTML抓取时的list_to_dataframe报错

这个错误的核心原因是:fetch_adjusted函数在处理某些日期时,返回的结果不是标准的数据框(比如页面解析失败返回了NULL,或者某天的表格结构和其他日期不一致)。而ldply要求所有函数返回的结果要么都是原子向量,要么都是数据框,类型不统一就会触发这个报错。

我给你调整了代码,添加了错误捕获机制,确保哪怕单个日期处理失败,整个批量任务也能继续,且始终返回结构一致的数据框:

library(plyr)
library(XML)

fetch_adjusted <- function(day) {
  # 把日期格式化为两位数字(比如1变成01),避免URL格式不规范
  day_str <- sprintf("%02d", day)
  fname <- paste0("standings201909", day_str, ".html")
  url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=2019-09-", day_str)
  
  # 捕获下载过程中的错误
  download_success <- tryCatch({
    download.file(url = url, destfile = fname)
    TRUE
  }, error = function(e) {
    message(paste("下载", day_str, "日数据失败:", e$message))
    FALSE
  })
  
  if (!download_success) {
    # 返回和正常结果结构一致的空数据框
    return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0)))
  }
  
  # 捕获HTML解析错误
  doc0 <- tryCatch(htmlParse(file = fname, encoding = "UTF-8"), error = function(e) {
    message(paste("解析", day_str, "日HTML失败:", e$message))
    NULL
  })
  
  if (is.null(doc0)) {
    return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0)))
  }
  
  doc1 <- xmlRoot(doc0)
  doc2 <- getNodeSet(doc1, "//table[@id='content']")
  
  if (length(doc2) == 0) {
    message(paste(day_str, "日未找到目标表格"))
    return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0)))
  }
  
  # 捕获表格读取错误
  standings <- tryCatch(readHTMLTable(doc2[[1]], header = TRUE, skip.rows = 1, stringsAsFactors = FALSE), error = function(e) {
    message(paste("读取", day_str, "日表格失败:", e$message))
    NULL
  })
  
  if (is.null(standings) || length(standings) == 0) {
    return(data.frame(matrix(ncol = ncol(standings_template), nrow = 0)))
  }
  
  standings <- standings[[1]]
  standings$day <- day_str
  standings
}

# 先获取一个正常的模板数据框,用来定义空数据框的列结构
standings_template <- fetch_adjusted(1)
# 批量处理9月1-29日的数据
Sept <- ldply(1:29, fetch_adjusted, .progress = "text")

代码说明:

  • 用sprintf格式化日期为两位数字,保证URL格式统一
  • 多层tryCatch捕获下载、解析、读取表格各环节的错误,避免单个日期失败导致整个任务中断
  • 定义standings_template确保空数据框和正常结果列数一致,让ldply能顺利合并
问题2:XLS文件被反复覆盖

这个问题很直观:你每次下载都用同一个文件名test.xls,新文件会直接覆盖旧文件,最后自然只剩最后一次的结果。另外你提到的“空表格”,大概率是最后一天的XLS本身为空,或者下载时没正确用二进制模式(wb是XLS文件的正确下载模式)。

给你两个解决方案,按需选择:

方案1:保存为多个独立的XLS文件(每个日期一个)

修改文件名,让每个文件包含对应的日期,这样就不会互相覆盖:

dates <- seq(as.Date("2019-09-01"), as.Date("2019-09-30"), by = 1)

fetch_adjusted <- function(date) {
  date_str <- as.character(date)
  url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=", date_str, "&otype=xls")
  # 生成带日期的唯一文件名
  destfile <- paste0("standings_", date_str, ".xls")
  
  # 捕获下载错误,用wb模式下载二进制文件
  tryCatch({
    download.file(url, destfile, mode = "wb")
    message(paste("成功下载:", destfile))
  }, error = function(e) {
    message(paste("下载", date_str, "失败:", e$message))
  })
}

# 批量下载所有日期的文件
lapply(dates, fetch_adjusted)

方案2:直接读取并合并成一个数据框(不保留单个文件)

如果不需要保存原始XLS文件,而是直接得到合并后的数据集,可以用临时文件来处理:

library(plyr)
library(readxl) # 需要先安装:install.packages("readxl")

dates <- seq(as.Date("2019-09-01"), as.Date("2019-09-30"), by = 1)

fetch_adjusted <- function(date) {
  date_str <- as.character(date)
  url <- paste0("https://legacy.baseballprospectus.com/standings/index.php?odate=", date_str, "&otype=xls")
  # 创建临时文件,下载后读取就删除
  temp_file <- tempfile(fileext = ".xls")
  
  # 捕获下载错误
  download_success <- tryCatch({
    download.file(url, temp_file, mode = "wb")
    TRUE
  }, error = function(e) {
    message(paste("下载", date_str, "失败:", e$message))
    FALSE
  })
  
  if (!download_success) {
    return(data.frame())
  }
  
  # 读取XLS文件,可根据实际表格结构调整skip、col_names等参数
  standings <- tryCatch(read_excel(temp_file), error = function(e) {
    message(paste("读取", date_str, "数据失败:", e$message))
    data.frame()
  })
  
  # 删除临时文件
  file.remove(temp_file)
  
  # 添加日期标识列
  standings$date <- date_str
  standings
}

# 批量处理并合并成一个大的数据框
combined_standings <- ldply(dates, fetch_adjusted, .progress = "text")

方案说明:

  • 方案1适合需要保留原始文件归档的场景,每个文件有唯一的日期命名
  • 方案2适合直接做数据分析的场景,用临时文件避免磁盘冗余,一步得到合并后的数据集

内容的提问来源于stack exchange,提问作者Abb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:17:47