如何用R程序化获取美国国家档案馆系列目录的可下载文件信息
Great question! Since archives.gov uses JavaScript to render content, straightforward rvest calls won't work for grabbing dynamic elements like download links directly from static HTML. Here's a step-by-step solution using R, with two approaches—one leveraging the site's hidden API (faster and more reliable) and another using RSelenium for full browser rendering if the API route doesn't pan out:
Step 1: Get All File Units in the Series
First, instead of scraping the series page, use the same API that powers the "Export" button to fetch all file unit entries directly as CSV. This avoids dealing with JS-rendered list items entirely.
- Open your target series page in a browser, open the developer tools (F12), click "Export", and look for the network request that loads the CSV. Copy that API URL.
- Use
httrandreadrto fetch and parse the CSV:
library(httr) library(readr) # Replace with your series' export API URL (from browser dev tools) series_export_url <- "https://catalog.archives.gov/api/v1/fileUnits?seriesId=18491489&export=csv" # Fetch and parse the CSV response <- GET(series_export_url) file_units <- read_csv(content(response, "text")) # Keep only the columns we need file_units_slim <- file_units[, c("naId", "title", "url")] head(file_units_slim)
Step 2: Fetch Download Links for Each File Unit
Approach 1: Use the File Unit API (Preferred)
Most file unit pages pull their download file data from a dedicated API endpoint. You can call this directly instead of loading the full page:
library(jsonlite) # Function to fetch download files for a single file unit ID get_download_files <- function(na_id) { api_url <- paste0("https://catalog.archives.gov/api/v1/fileUnits/", na_id, "/files") response <- GET(api_url) # Parse JSON response file_data <- fromJSON(content(response, "text")) # Extract file titles and download URLs if (!is.null(file_data$data)) { return(data.frame( file_unit_title = file_units$title[file_units$naId == na_id], file_title = file_data$data$title, file_url = file_data$data$downloadUrl, stringsAsFactors = FALSE )) } else { return(NULL) } } # Loop through all file units to collect download links all_downloads <- lapply(file_units_slim$naId, get_download_files) all_downloads <- do.call(rbind, all_downloads) # View results head(all_downloads)
Approach 2: Use RSelenium for Full Browser Rendering
If the API route doesn't work (e.g., the site changes its API structure), use RSelenium to simulate a browser and interact with the JS-rendered pages:
library(RSelenium) library(rvest) # Start a Chrome browser instance (ensure ChromeDriver is installed and accessible) driver <- rsDriver(browser = "chrome", verbose = FALSE) remDr <- driver[["client"]] # Initialize list to store results all_downloads <- list() for (i in seq_len(nrow(file_units_slim))) { unit_url <- file_units_slim$url[i] unit_title <- file_units_slim$title[i] # Navigate to the file unit page remDr$navigate(unit_url) # Wait for JS to render content (adjust sleep time if needed) Sys.sleep(3) # Extract page HTML and parse with rvest page_html <- remDr$getPageSource()[[1]] %>% read_html() # Grab download links and titles (adjust CSS selector to match the site's structure) download_links <- page_html %>% html_elements(".download a") %>% html_attr("href") download_titles <- page_html %>% html_elements(".download a") %>% html_text(trim = TRUE) # Add to results if there are downloads if (length(download_links) > 0) { all_downloads[[i]] <- data.frame( file_unit_title = unit_title, file_title = download_titles, file_url = download_links, stringsAsFactors = FALSE ) } } # Combine results and clean up all_downloads <- do.call(rbind, all_downloads) remDr$close() driver$server$stop() # View results head(all_downloads)
Notes
- Always check the site's
robots.txt(https://catalog.archives.gov/robots.txt) to ensure your scraping is allowed. Add delays between requests to avoid overwhelming the server. - If using
RSelenium, you'll need to have Chrome (or another browser) and the corresponding driver installed. Usewdman::chrome_driver()to manage ChromeDriver versions easily.
内容的提问来源于stack exchange,提问作者Anthony Damico

