如何下载含数字开头年份文件的文件夹?rvest下载Zip变量为空求助
First, let's tackle why your zips variable is empty—you've got a small typo in your code, plus a couple of tweaks to make it more reliable:
1. Correct the Variable Name & Selector
You assigned the parsed HTML to rh, but then tried to use an undefined pg variable in html_nodes(pg, ...)—that's the main reason nothing was found. Let's fix that, plus adjust the selector to target all ZIP files directly (more robust than filtering by the year prefix):
library(rvest) library(purrr) # For easy iteration over files # Base URL of the folder base_url <- "https://danepubliczne.imgw.pl/data/dane_pomiarowo_obserwacyjne/dane_meteorologiczne/dobowe/synop/2011/" # Parse the webpage page <- read_html(base_url) # Extract all links pointing to ZIP files zips <- page %>% html_elements("a[href$='.zip']") %>% # Target links ending with .zip html_attr("href")
Now zips should contain all the ZIP filenames from the page.
2. Download All ZIP Files to a Local Folder
Once you have the list of ZIPs, combine each filename with the base URL and download them to an organized local folder:
# Create a local folder to store downloads (if it doesn't exist) local_folder <- "imgw_2011_synop_data" if (!dir.exists(local_folder)) dir.create(local_folder) # Download each file (add a small delay to avoid overwhelming the server) walk(zips, function(zip_file) { file_url <- paste0(base_url, zip_file) local_path <- file.path(local_folder, zip_file) download.file(file_url, local_path, mode = "wb") # mode="wb" for binary files Sys.sleep(0.5) # Optional polite delay to prevent hitting server limits })
Web servers typically don't support direct "folder downloads" via browser or R, but the method above effectively replicates that by:
- Scraping all relevant file links from the folder's webpage
- Downloading each file individually to a local folder
If you're comfortable with command-line tools, you can also use wget (available on most operating systems) to do this in one line:
wget -r -np -A.zip https://danepubliczne.imgw.pl/data/dane_pomiarowo_obserwacyjne/dane_meteorologiczne/dobowe/synop/2011/
-r: Enable recursive download-np: Don't traverse to parent directories-A.zip: Only download files ending with.zip
内容的提问来源于stack exchange,提问作者Tom

