You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R批量下载指定URL的CSV文件并提取指定行时,下载目录为空的问题求助

使用R批量下载指定URL的CSV文件并提取指定行时,下载目录为空的问题求助

各位好,我最近尝试用R脚本从MISO Energy的网页批量下载历史MCP的CSV文件,然后提取每个文件里第3到13行的数据,但运行代码后,桌面上虽然生成了downloaded_files文件夹,但里面是空的,完全没有下载到任何文件。我不确定是代码哪里出了问题,还是目标网页本身有什么限制需要处理,想请大家帮忙看看。

我的代码如下:

library(rvest)

url <- "https://www.misoenergy.org/markets-and-operations/real-time--market-data/market-reports/#nt=%2FMarketReportType%3AHistorical%20MCP%2FMarketReportName%3AASM%20Real-Time%20Final%20Market%20MCPs%20(csv)&t=10&p=0&s=MarketReportPublished&sd=desc"

# Scrape the webpage to extract the URLs
page <- read_html(url)
file_links <- page %>% html_nodes("a[href$='.csv']") %>% html_attr("href")

# Create a directory to store the downloaded files
dir.create("downloaded_files")

# Define the range of rows to extract from each file
start_row <- 3
end_row <- 13

# Loop over each file URL and download the file
for (file_link in file_links) {
  filename <- basename(file_link)
  file_path <- paste0("downloaded_files/", filename)
  
  # Download the file
  download.file(file_link, destfile = file_path)
  
  # Read the downloaded file
  data <- read.csv(file_path, skip = start_row - 1, nrows = end_row - start_row + 1)
  
  # Do something with the data, e.g., print the extracted rows
  print(data)
}

我已经确认dir.create生成的文件夹路径是对的,但就是没有文件被下载进来。目前自己怀疑几个方向:

  • 是不是网页的CSV链接是动态加载的,rvest直接爬取静态HTML拿不到真实的下载链接?
  • 下载的时候是不是需要给download.file指定method参数(比如用"curl"或者"wget")?
  • 目标网页有没有反爬机制,需要添加请求头才能正常获取链接?

希望有经验的朋友能帮我排查一下问题,谢谢!

备注:内容来源于stack exchange,提问作者user22153287

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 11:04:49