R循环下载JSON数据时遇file与open.connection报错求助
Fixing Connection Errors in Your KVK JSON Download Script
Hey there, let's work through those file(con, "r") and open.connection(con, "rb") errors you're hitting. These almost always stem from unhandled request failures, aggressive scraping, or outdated package usage. Here's how to fix it step by step:
Common Causes of Your Errors
- No error handling: When the server returns a 404, 503, or times out,
RJSONIO::fromJSONcrashes instead of recovering. - Misplaced timeout setting: Your
options(timeout)is inside the loop, so it might not take effect for every request. - Unrestricted loop range: You're looping up to 99999, but most of those pages likely don't exist—servers block repeated invalid requests.
- Outdated package:
RJSONIOis no longer actively maintained; modern alternatives likejsonlitehandle errors far better.
Fixed & Improved Script
First, install and use jsonlite (it's more reliable), then add error handling, rate limiting, and smarter loop controls:
# Install jsonlite if you haven't already if (!require(jsonlite)) { install.packages("jsonlite") library(jsonlite) } # Base URL for KVK requests baseurl <- "http://zoeken.kvk.nl/Address.ashx?site=handelsregister&partialfields=&q=010" pages <- list() # Set a reasonable timeout (300 seconds = 5 mins) once, outside the loop options(timeout = 300) # Configure retry and rate-limiting settings max_retries <- 3 # Max attempts per page request_delay <- 2 # Seconds to wait between retries batch_delay <- 5 # Seconds to wait after every 10 requests # First, test to find valid page ranges (adjust start/end based on your testing) start_page <- 10000 end_page <- 12000 # Replace with actual valid upper limit after testing for(i in start_page:end_page){ message("Retrieving page ", i) success <- FALSE retries <- 0 # Retry logic for temporary failures while(!success && retries < max_retries){ tryCatch({ full_url <- paste0(baseurl, i) mydata <- fromJSON(full_url, flatten = TRUE) # Only save data if results exist if(!is.null(mydata$resultatenHR) && length(mydata$resultatenHR) > 0){ pages[[as.character(i)]] <- mydata$resultatenHR # Use page number as list key for clarity success <- TRUE } else { message("Page ", i, " has no results—skipping.") success <- TRUE } }, error = function(e){ retries <<- retries + 1 message("Failed attempt ", retries, " for page ", i, ": ", e$message) Sys.sleep(request_delay) }) } # Take a break every 10 requests to avoid triggering anti-scraping rules if(i %% 10 == 0){ message("Taking a short break to respect server limits...") Sys.sleep(batch_delay) } # Optional: Save progress periodically to avoid data loss if(i %% 100 == 0){ saveRDS(pages, paste0("kvk_progress_", i, ".rds")) message("Saved progress to kvk_progress_", i, ".rds") } } # Save final dataset saveRDS(pages, "kvk_full_dataset.rds") message("Download complete! Data saved to kvk_full_dataset.rds")
Key Improvements Explained
- Error handling with
tryCatch: Catches connection failures and retries instead of crashing the loop. - Rate limiting: Adds delays between requests to avoid being blocked by the KVK server.
- Progress saving: Periodically saves your data to a file, so you don't lose everything if the script crashes.
- Valid page range: Forces you to test first and only request pages that actually exist, reducing invalid requests.
- Modern package:
jsonliteis actively maintained and handles JSON parsing more robustly thanRJSONIO.
Additional Tips
- Check
robots.txt: Always verifyhttp://zoeken.kvk.nl/robots.txtto make sure scraping this endpoint is allowed (respect website terms of service!). - Adjust page range: Test a few high-numbered pages first (e.g., 10000, 11000, 12000) to find the actual upper limit of valid results—don't waste time on 99999 pages if only 15000 exist.
- Avoid parallel requests unless necessary: While parallel scraping can speed things up, it's more likely to get you blocked. Stick to sequential requests with delays unless you're sure the server allows it.
内容的提问来源于stack exchange,提问作者RobertHaa
相关产品推荐
相关产品推荐

