技术求助:使用R抓取多页面表格数据(附目标网站)
Hey there! Let's walk through how to pull all the data from that multi-page FDIC failed banks table into your R environment. This approach uses common web scraping packages and handles pagination seamlessly.
Step 1: Install & Load Required Packages
First, make sure you have the necessary packages installed. We'll use rvest for web scraping, dplyr for data manipulation, and purrr for iterating over pages:
# Install packages if you haven't already install.packages(c("rvest", "dplyr", "purrr")) # Load the packages library(rvest) library(dplyr) library(purrr)
Step 2: Understand the Pagination Structure
The FDIC bank list uses query parameters for pagination. Each page is accessible via https://www.fdic.gov/bank/individual/failed/banklist.html?page=X where X is the page number. We'll first figure out how many total pages there are, then scrape each one.
Step 3: Create a Function to Scrape a Single Page
Let's write a helper function that takes a page number and returns the data from that page as a data frame. We'll add a tryCatch to handle any potential errors (like a page that doesn't exist):
scrape_fdic_page <- function(page_num) { # Build the URL for the target page url <- paste0("https://www.fdic.gov/bank/individual/failed/banklist.html?page=", page_num) # Try to scrape the page, return an empty data frame if it fails tryCatch({ # Read the HTML content page <- read_html(url) # Extract the table (the first table on the page is our target) table_data <- page %>% html_element("table") %>% html_table(header = TRUE) # Return the data frame return(table_data) }, error = function(e) { message(paste("Failed to scrape page", page_num, ":", e$message)) return(data.frame()) }) }
Step 4: Scrape All Pages & Combine Data
First, let's find the total number of pages. We can get this from the pagination navigation on the first page:
# Scrape the first page to get pagination info first_page <- read_html("https://www.fdic.gov/bank/individual/failed/banklist.html") # Extract the total number of pages from the pagination text total_pages <- first_page %>% html_element(".pagination") %>% html_text() %>% stringr::str_extract("of (\\d+)") %>% stringr::str_remove("of ") %>% as.integer()
Now, iterate over all pages, scrape each one, and combine the results into a single data frame:
# Scrape all pages using map_dfr to automatically bind rows all_failed_banks <- map_dfr(1:total_pages, scrape_fdic_page) # Check the result head(all_failed_banks) nrow(all_failed_banks)
Bonus: Clean Up the Data (Optional)
The scraped data might have some formatting quirks. For example, you might want to convert dates to proper date types:
all_failed_banks_cleaned <- all_failed_banks %>% mutate( `Closing Date` = as.Date(`Closing Date`, "%B %d, %Y"), `Fund` = as.numeric(`Fund`) ) # View the cleaned data glimpse(all_failed_banks_cleaned)
That's it! You should now have all the failed bank data from every page loaded into your R environment.
内容的提问来源于stack exchange,提问作者RSOLK

