使用R批量下载指定URL的CSV文件并提取指定行时,下载目录为空的问题求助
使用R批量下载指定URL的CSV文件并提取指定行时,下载目录为空的问题求助
各位好,我最近尝试用R脚本从MISO Energy的网页批量下载历史MCP的CSV文件,然后提取每个文件里第3到13行的数据,但运行代码后,桌面上虽然生成了downloaded_files文件夹,但里面是空的,完全没有下载到任何文件。我不确定是代码哪里出了问题,还是目标网页本身有什么限制需要处理,想请大家帮忙看看。
我的代码如下:
library(rvest) url <- "https://www.misoenergy.org/markets-and-operations/real-time--market-data/market-reports/#nt=%2FMarketReportType%3AHistorical%20MCP%2FMarketReportName%3AASM%20Real-Time%20Final%20Market%20MCPs%20(csv)&t=10&p=0&s=MarketReportPublished&sd=desc" # Scrape the webpage to extract the URLs page <- read_html(url) file_links <- page %>% html_nodes("a[href$='.csv']") %>% html_attr("href") # Create a directory to store the downloaded files dir.create("downloaded_files") # Define the range of rows to extract from each file start_row <- 3 end_row <- 13 # Loop over each file URL and download the file for (file_link in file_links) { filename <- basename(file_link) file_path <- paste0("downloaded_files/", filename) # Download the file download.file(file_link, destfile = file_path) # Read the downloaded file data <- read.csv(file_path, skip = start_row - 1, nrows = end_row - start_row + 1) # Do something with the data, e.g., print the extracted rows print(data) }
我已经确认dir.create生成的文件夹路径是对的,但就是没有文件被下载进来。目前自己怀疑几个方向:
- 是不是网页的CSV链接是动态加载的,
rvest直接爬取静态HTML拿不到真实的下载链接? - 下载的时候是不是需要给
download.file指定method参数(比如用"curl"或者"wget")? - 目标网页有没有反爬机制,需要添加请求头才能正常获取链接?
希望有经验的朋友能帮我排查一下问题,谢谢!
备注:内容来源于stack exchange,提问作者user22153287
相关产品推荐
相关产品推荐

