You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest读取HTML表格时出现超时卡顿问题,如何稳定爬取kursdollar.org各银行美元汇率表格?

Problem: Unstable Web Scraping of Exchange Rate Tables in R

I'm trying to scrape US dollar exchange rate tables from various bank pages on kursdollar.org using the R code below. However, the results are inconsistent: sometimes it works quickly, but other times it gets stuck on a link (the RCurl timeout setting doesn't seem to take effect) and throws this error:

Error in url(link_kurs[v], "rb") : cannot open the connection
In addition: Warning message:
In url(link_kurs[v], "rb") : InternetOpenUrl failed: 'The operation timed out'

I need a way to reliably read all the tables, even if it's a bit slower. Here's the code I've been testing:

library(stringr)
library(tidyverse)
library(rvest)
library(httr)
library(RCurl)
curlSetOpt(timeout = 200)

kurs_bi <- "https://kursdollar.org/bank/bi.php"
kurs_mandiri <- "https://kursdollar.org/bank/mandiri.php"
kurs_bca <- "https://kursdollar.org/bank/bca.php"
kurs_bni <- "https://kursdollar.org/bank/bni.php"
kurs_hsbc <- "https://kursdollar.org/bank/hsbc.php"
kurs_panin <- "https://kursdollar.org/bank/panin.php"
kurs_cimb <- "https://kursdollar.org/bank/cimb.php"
kurs_ocbc <- "https://kursdollar.org/bank/ocbc.php"
kurs_bri <- "https://kursdollar.org/bank/bri.php"
kurs_uob <- "https://kursdollar.org/bank/uob.php"
kurs_maybank <- 'https://kursdollar.org/bank/maybank.php'
kurs_permata <- "https://kursdollar.org/bank/permata.php"
kurs_mega <- "https://kursdollar.org/bank/mega.php"
kurs_danamon <- "https://kursdollar.org/bank/danamon.php"
kurs_btn <- "https://kursdollar.org/bank/btn.php"
kurs_mayapada <- "https://kursdollar.org/bank/mayapada.php"
kurs_muamalat <- "https://kursdollar.org/bank/muamalat.php"
kurs_bukopin <- "https://kursdollar.org/bank/bukopin.php"

link_kurs <- c(kurs_bi, kurs_mandiri, kurs_bca, kurs_bni, kurs_hsbc, kurs_panin, kurs_cimb, kurs_ocbc, kurs_bri, kurs_uob, kurs_maybank, kurs_permata, kurs_mega, kurs_danamon, kurs_btn, kurs_mayapada, kurs_muamalat, kurs_bukopin)

for(v in 1:length(link_kurs)){
  writeLines(paste0(v,') Read Table on ', link_kurs[v]))
  open_url <- url(link_kurs[v], "rb")
  extract_df <- read_html(open_url)
  close(open_url)
  extract_df <- extract_df %>% html_nodes("table") %>% html_table(fill = T) %>% as.data.frame()
  writeLines("Test Read Success!")
}

Solution: Robust Web Scraping with Timeouts and Retries

The main issues with your current code are:

  • The url() function doesn't respect RCurl's timeout settings (they're separate libraries with their own connection handling).
  • No retry mechanism for temporary network glitches or server delays.
  • No delays between requests, which might trigger the site's rate-limiting defenses.

Here's an improved version of your code that fixes these problems and ensures stable scraping:

library(tidyverse)
library(rvest)
library(httr)
library(retry) # Install first with install.packages("retry")

# Define all bank URLs in a cleaner vector
link_kurs <- c(
  "https://kursdollar.org/bank/bi.php",
  "https://kursdollar.org/bank/mandiri.php",
  "https://kursdollar.org/bank/bca.php",
  "https://kursdollar.org/bank/bni.php",
  "https://kursdollar.org/bank/hsbc.php",
  "https://kursdollar.org/bank/panin.php",
  "https://kursdollar.org/bank/cimb.php",
  "https://kursdollar.org/bank/ocbc.php",
  "https://kursdollar.org/bank/bri.php",
  "https://kursdollar.org/bank/uob.php",
  "https://kursdollar.org/bank/maybank.php",
  "https://kursdollar.org/bank/permata.php",
  "https://kursdollar.org/bank/mega.php",
  "https://kursdollar.org/bank/danamon.php",
  "https://kursdollar.org/bank/btn.php",
  "https://kursdollar.org/bank/mayapada.php",
  "https://kursdollar.org/bank/muamalat.php",
  "https://kursdollar.org/bank/bukopin.php"
)

# Define a safe, retry-enabled scraping function
safe_scrape_table <- function(url) {
  # Retry failed requests up to 3 times, with 2-second delays between attempts
  result <- retry({
    # Use httr::GET with explicit 10-second timeout (works consistently)
    response <- GET(url, timeout(10), 
                    user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"))
    
    # Stop if server returns an error (e.g., 404, 500)
    stop_for_status(response)
    
    # Parse HTML and extract tables
    html <- content(response, "text") %>% read_html()
    tables <- html %>% html_nodes("table") %>% html_table(fill = TRUE)
    
    if (length(tables) == 0) {
      warning(paste("No tables found on:", url))
      return(NULL)
    }
    
    # Return the first table (adjust to tables[[n]] if you need a specific one)
    as.data.frame(tables[[1]])
  },
  max_tries = 3,
  delay = 2,
  quiet = FALSE
  )
  
  result
}

# Scrape all URLs with progress tracking
scraped_data <- map(link_kurs, function(url) {
  message(paste("Processing:", url))
  df <- safe_scrape_table(url)
  
  # Add a column to identify the source URL (useful for debugging)
  if (!is.null(df)) {
    df$source_bank <- str_extract(url, "(?<=bank/).*(?=\\.php)")
  }
  
  df
})

# Combine all successful results into a single dataframe (optional)
combined_exchange_rates <- bind_rows(scraped_data)

Key Improvements Explained:

  1. Reliable Timeouts: We use httr::GET() with timeout(10) to enforce a hard timeout per request—this works consistently, unlike RCurl settings paired with the base url() function.
  2. Retry Mechanism: The retry package handles temporary network issues by reattempting failed requests up to 3 times, with a 2-second delay between tries.
  3. Rate Limiting: The delay between retries (plus natural processing time) reduces the chance of being blocked by the site's anti-scraping measures.
  4. Error Handling: stop_for_status() checks for server errors and stops gracefully, while warnings alert you if a page has no tables.
  5. Cleaner Code: We use purrr::map() instead of a for loop for more idiomatic R, and add a source_bank column to track which table came from which bank.

Additional Tips:

  • If you still hit timeouts, increase the timeout value (e.g., timeout(15)) or the number of retries (max_tries = 5).
  • Always check the site's robots.txt to confirm scraping is allowed.

内容的提问来源于stack exchange,提问作者Jovan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:17:36