You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R程序化获取美国国家档案馆系列目录的可下载文件信息

Great question! Since archives.gov uses JavaScript to render content, straightforward rvest calls won't work for grabbing dynamic elements like download links directly from static HTML. Here's a step-by-step solution using R, with two approaches—one leveraging the site's hidden API (faster and more reliable) and another using RSelenium for full browser rendering if the API route doesn't pan out:

Step 1: Get All File Units in the Series

First, instead of scraping the series page, use the same API that powers the "Export" button to fetch all file unit entries directly as CSV. This avoids dealing with JS-rendered list items entirely.

  1. Open your target series page in a browser, open the developer tools (F12), click "Export", and look for the network request that loads the CSV. Copy that API URL.
  2. Use httr and readr to fetch and parse the CSV:
library(httr)
library(readr)

# Replace with your series' export API URL (from browser dev tools)
series_export_url <- "https://catalog.archives.gov/api/v1/fileUnits?seriesId=18491489&export=csv"

# Fetch and parse the CSV
response <- GET(series_export_url)
file_units <- read_csv(content(response, "text"))

# Keep only the columns we need
file_units_slim <- file_units[, c("naId", "title", "url")]
head(file_units_slim)

Approach 1: Use the File Unit API (Preferred)

Most file unit pages pull their download file data from a dedicated API endpoint. You can call this directly instead of loading the full page:

library(jsonlite)

# Function to fetch download files for a single file unit ID
get_download_files <- function(na_id) {
  api_url <- paste0("https://catalog.archives.gov/api/v1/fileUnits/", na_id, "/files")
  response <- GET(api_url)
  
  # Parse JSON response
  file_data <- fromJSON(content(response, "text"))
  
  # Extract file titles and download URLs
  if (!is.null(file_data$data)) {
    return(data.frame(
      file_unit_title = file_units$title[file_units$naId == na_id],
      file_title = file_data$data$title,
      file_url = file_data$data$downloadUrl,
      stringsAsFactors = FALSE
    ))
  } else {
    return(NULL)
  }
}

# Loop through all file units to collect download links
all_downloads <- lapply(file_units_slim$naId, get_download_files)
all_downloads <- do.call(rbind, all_downloads)

# View results
head(all_downloads)

Approach 2: Use RSelenium for Full Browser Rendering

If the API route doesn't work (e.g., the site changes its API structure), use RSelenium to simulate a browser and interact with the JS-rendered pages:

library(RSelenium)
library(rvest)

# Start a Chrome browser instance (ensure ChromeDriver is installed and accessible)
driver <- rsDriver(browser = "chrome", verbose = FALSE)
remDr <- driver[["client"]]

# Initialize list to store results
all_downloads <- list()

for (i in seq_len(nrow(file_units_slim))) {
  unit_url <- file_units_slim$url[i]
  unit_title <- file_units_slim$title[i]
  
  # Navigate to the file unit page
  remDr$navigate(unit_url)
  
  # Wait for JS to render content (adjust sleep time if needed)
  Sys.sleep(3)
  
  # Extract page HTML and parse with rvest
  page_html <- remDr$getPageSource()[[1]] %>% read_html()
  
  # Grab download links and titles (adjust CSS selector to match the site's structure)
  download_links <- page_html %>% 
    html_elements(".download a") %>% 
    html_attr("href")
  
  download_titles <- page_html %>% 
    html_elements(".download a") %>% 
    html_text(trim = TRUE)
  
  # Add to results if there are downloads
  if (length(download_links) > 0) {
    all_downloads[[i]] <- data.frame(
      file_unit_title = unit_title,
      file_title = download_titles,
      file_url = download_links,
      stringsAsFactors = FALSE
    )
  }
}

# Combine results and clean up
all_downloads <- do.call(rbind, all_downloads)
remDr$close()
driver$server$stop()

# View results
head(all_downloads)

Notes

  • Always check the site's robots.txt (https://catalog.archives.gov/robots.txt) to ensure your scraping is allowed. Add delays between requests to avoid overwhelming the server.
  • If using RSelenium, you'll need to have Chrome (or another browser) and the corresponding driver installed. Use wdman::chrome_driver() to manage ChromeDriver versions easily.

内容的提问来源于stack exchange,提问作者Anthony Damico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:28:53