使用rvest读取HTML表格时出现超时卡顿问题,如何稳定爬取kursdollar.org各银行美元汇率表格?
Problem: Unstable Web Scraping of Exchange Rate Tables in R
I'm trying to scrape US dollar exchange rate tables from various bank pages on kursdollar.org using the R code below. However, the results are inconsistent: sometimes it works quickly, but other times it gets stuck on a link (the RCurl timeout setting doesn't seem to take effect) and throws this error:
Error in url(link_kurs[v], "rb") : cannot open the connection In addition: Warning message: In url(link_kurs[v], "rb") : InternetOpenUrl failed: 'The operation timed out'
I need a way to reliably read all the tables, even if it's a bit slower. Here's the code I've been testing:
library(stringr) library(tidyverse) library(rvest) library(httr) library(RCurl) curlSetOpt(timeout = 200) kurs_bi <- "https://kursdollar.org/bank/bi.php" kurs_mandiri <- "https://kursdollar.org/bank/mandiri.php" kurs_bca <- "https://kursdollar.org/bank/bca.php" kurs_bni <- "https://kursdollar.org/bank/bni.php" kurs_hsbc <- "https://kursdollar.org/bank/hsbc.php" kurs_panin <- "https://kursdollar.org/bank/panin.php" kurs_cimb <- "https://kursdollar.org/bank/cimb.php" kurs_ocbc <- "https://kursdollar.org/bank/ocbc.php" kurs_bri <- "https://kursdollar.org/bank/bri.php" kurs_uob <- "https://kursdollar.org/bank/uob.php" kurs_maybank <- 'https://kursdollar.org/bank/maybank.php' kurs_permata <- "https://kursdollar.org/bank/permata.php" kurs_mega <- "https://kursdollar.org/bank/mega.php" kurs_danamon <- "https://kursdollar.org/bank/danamon.php" kurs_btn <- "https://kursdollar.org/bank/btn.php" kurs_mayapada <- "https://kursdollar.org/bank/mayapada.php" kurs_muamalat <- "https://kursdollar.org/bank/muamalat.php" kurs_bukopin <- "https://kursdollar.org/bank/bukopin.php" link_kurs <- c(kurs_bi, kurs_mandiri, kurs_bca, kurs_bni, kurs_hsbc, kurs_panin, kurs_cimb, kurs_ocbc, kurs_bri, kurs_uob, kurs_maybank, kurs_permata, kurs_mega, kurs_danamon, kurs_btn, kurs_mayapada, kurs_muamalat, kurs_bukopin) for(v in 1:length(link_kurs)){ writeLines(paste0(v,') Read Table on ', link_kurs[v])) open_url <- url(link_kurs[v], "rb") extract_df <- read_html(open_url) close(open_url) extract_df <- extract_df %>% html_nodes("table") %>% html_table(fill = T) %>% as.data.frame() writeLines("Test Read Success!") }
Solution: Robust Web Scraping with Timeouts and Retries
The main issues with your current code are:
- The
url()function doesn't respectRCurl's timeout settings (they're separate libraries with their own connection handling). - No retry mechanism for temporary network glitches or server delays.
- No delays between requests, which might trigger the site's rate-limiting defenses.
Here's an improved version of your code that fixes these problems and ensures stable scraping:
library(tidyverse) library(rvest) library(httr) library(retry) # Install first with install.packages("retry") # Define all bank URLs in a cleaner vector link_kurs <- c( "https://kursdollar.org/bank/bi.php", "https://kursdollar.org/bank/mandiri.php", "https://kursdollar.org/bank/bca.php", "https://kursdollar.org/bank/bni.php", "https://kursdollar.org/bank/hsbc.php", "https://kursdollar.org/bank/panin.php", "https://kursdollar.org/bank/cimb.php", "https://kursdollar.org/bank/ocbc.php", "https://kursdollar.org/bank/bri.php", "https://kursdollar.org/bank/uob.php", "https://kursdollar.org/bank/maybank.php", "https://kursdollar.org/bank/permata.php", "https://kursdollar.org/bank/mega.php", "https://kursdollar.org/bank/danamon.php", "https://kursdollar.org/bank/btn.php", "https://kursdollar.org/bank/mayapada.php", "https://kursdollar.org/bank/muamalat.php", "https://kursdollar.org/bank/bukopin.php" ) # Define a safe, retry-enabled scraping function safe_scrape_table <- function(url) { # Retry failed requests up to 3 times, with 2-second delays between attempts result <- retry({ # Use httr::GET with explicit 10-second timeout (works consistently) response <- GET(url, timeout(10), user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")) # Stop if server returns an error (e.g., 404, 500) stop_for_status(response) # Parse HTML and extract tables html <- content(response, "text") %>% read_html() tables <- html %>% html_nodes("table") %>% html_table(fill = TRUE) if (length(tables) == 0) { warning(paste("No tables found on:", url)) return(NULL) } # Return the first table (adjust to tables[[n]] if you need a specific one) as.data.frame(tables[[1]]) }, max_tries = 3, delay = 2, quiet = FALSE ) result } # Scrape all URLs with progress tracking scraped_data <- map(link_kurs, function(url) { message(paste("Processing:", url)) df <- safe_scrape_table(url) # Add a column to identify the source URL (useful for debugging) if (!is.null(df)) { df$source_bank <- str_extract(url, "(?<=bank/).*(?=\\.php)") } df }) # Combine all successful results into a single dataframe (optional) combined_exchange_rates <- bind_rows(scraped_data)
Key Improvements Explained:
- Reliable Timeouts: We use
httr::GET()withtimeout(10)to enforce a hard timeout per request—this works consistently, unlikeRCurlsettings paired with the baseurl()function. - Retry Mechanism: The
retrypackage handles temporary network issues by reattempting failed requests up to 3 times, with a 2-second delay between tries. - Rate Limiting: The delay between retries (plus natural processing time) reduces the chance of being blocked by the site's anti-scraping measures.
- Error Handling:
stop_for_status()checks for server errors and stops gracefully, while warnings alert you if a page has no tables. - Cleaner Code: We use
purrr::map()instead of a for loop for more idiomatic R, and add asource_bankcolumn to track which table came from which bank.
Additional Tips:
- If you still hit timeouts, increase the
timeoutvalue (e.g.,timeout(15)) or the number of retries (max_tries = 5). - Always check the site's
robots.txtto confirm scraping is allowed.
内容的提问来源于stack exchange,提问作者Jovan
相关产品推荐
相关产品推荐

