You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R循环下载JSON数据时遇file与open.connection报错求助

Fixing Connection Errors in Your KVK JSON Download Script

Hey there, let's work through those file(con, "r") and open.connection(con, "rb") errors you're hitting. These almost always stem from unhandled request failures, aggressive scraping, or outdated package usage. Here's how to fix it step by step:

Common Causes of Your Errors

  • No error handling: When the server returns a 404, 503, or times out, RJSONIO::fromJSON crashes instead of recovering.
  • Misplaced timeout setting: Your options(timeout) is inside the loop, so it might not take effect for every request.
  • Unrestricted loop range: You're looping up to 99999, but most of those pages likely don't exist—servers block repeated invalid requests.
  • Outdated package: RJSONIO is no longer actively maintained; modern alternatives like jsonlite handle errors far better.

Fixed & Improved Script

First, install and use jsonlite (it's more reliable), then add error handling, rate limiting, and smarter loop controls:

# Install jsonlite if you haven't already
if (!require(jsonlite)) {
  install.packages("jsonlite")
  library(jsonlite)
}

# Base URL for KVK requests
baseurl <- "http://zoeken.kvk.nl/Address.ashx?site=handelsregister&partialfields=&q=010"
pages <- list()

# Set a reasonable timeout (300 seconds = 5 mins) once, outside the loop
options(timeout = 300)

# Configure retry and rate-limiting settings
max_retries <- 3  # Max attempts per page
request_delay <- 2  # Seconds to wait between retries
batch_delay <- 5  # Seconds to wait after every 10 requests

# First, test to find valid page ranges (adjust start/end based on your testing)
start_page <- 10000
end_page <- 12000  # Replace with actual valid upper limit after testing

for(i in start_page:end_page){
  message("Retrieving page ", i)
  success <- FALSE
  retries <- 0
  
  # Retry logic for temporary failures
  while(!success && retries < max_retries){
    tryCatch({
      full_url <- paste0(baseurl, i)
      mydata <- fromJSON(full_url, flatten = TRUE)
      
      # Only save data if results exist
      if(!is.null(mydata$resultatenHR) && length(mydata$resultatenHR) > 0){
        pages[[as.character(i)]] <- mydata$resultatenHR  # Use page number as list key for clarity
        success <- TRUE
      } else {
        message("Page ", i, " has no results—skipping.")
        success <- TRUE
      }
    }, error = function(e){
      retries <<- retries + 1
      message("Failed attempt ", retries, " for page ", i, ": ", e$message)
      Sys.sleep(request_delay)
    })
  }
  
  # Take a break every 10 requests to avoid triggering anti-scraping rules
  if(i %% 10 == 0){
    message("Taking a short break to respect server limits...")
    Sys.sleep(batch_delay)
  }
  
  # Optional: Save progress periodically to avoid data loss
  if(i %% 100 == 0){
    saveRDS(pages, paste0("kvk_progress_", i, ".rds"))
    message("Saved progress to kvk_progress_", i, ".rds")
  }
}

# Save final dataset
saveRDS(pages, "kvk_full_dataset.rds")
message("Download complete! Data saved to kvk_full_dataset.rds")

Key Improvements Explained

  • Error handling with tryCatch: Catches connection failures and retries instead of crashing the loop.
  • Rate limiting: Adds delays between requests to avoid being blocked by the KVK server.
  • Progress saving: Periodically saves your data to a file, so you don't lose everything if the script crashes.
  • Valid page range: Forces you to test first and only request pages that actually exist, reducing invalid requests.
  • Modern package: jsonlite is actively maintained and handles JSON parsing more robustly than RJSONIO.

Additional Tips

  • Check robots.txt: Always verify http://zoeken.kvk.nl/robots.txt to make sure scraping this endpoint is allowed (respect website terms of service!).
  • Adjust page range: Test a few high-numbered pages first (e.g., 10000, 11000, 12000) to find the actual upper limit of valid results—don't waste time on 99999 pages if only 15000 exist.
  • Avoid parallel requests unless necessary: While parallel scraping can speed things up, it's more likely to get you blocked. Stick to sequential requests with delays unless you're sure the server allows it.

内容的提问来源于stack exchange,提问作者RobertHaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:25:16