如何从URL列表读取JSON并整理为DataFrame?R语言IP地理编码求助
Hey there, let's work through your ipstack batch geocoding issue step by step. I'll break down the problems you ran into and give you a robust solution that handles edge cases like failed requests and missing data.
First, let's fix the core errors you encountered
Why
read_json(x)threw an error
Theread_json()function fromjsonliteonly accepts a single URL/file path as input, not a vector of URLs. You need to iterate over each URL individually to fetch data.Why your initial loop had row mismatch errors
When some IP requests fail (invalid IP, API rate limiting, network blips),read_json()might return an empty list, which converts to a 0-row data frame. Trying torbind()that with a 1-row data frame causes the "differing number of rows" error.Why the
zip is NULLerror happened
Some IP addresses don't have a zip code associated with them, so the API returnsNULLfor that field. Tibbles don't allow NULL columns—you need to convert those NULLs toNAinstead.
Robust Solution with Error Handling & Rate Limiting
Here's a refined function that handles all these issues, plus adds rate limiting to avoid getting blocked by ipstack (critical for 1200 requests):
First, load the required packages:
library(httr) library(jsonlite) library(purrr) library(dplyr) library(tibble)
Then define a robust IP lookup function with error catching:
ip_locate <- function(ip_address, access_key) { # Build the request URL request_url <- paste0("http://api.ipstack.com/", ip_address, "?access_key=", access_key) # Use tryCatch to handle failed requests gracefully tryCatch({ # Send the request and check for HTTP errors response <- httr::GET(request_url) httr::stop_for_status(response) # Throws error if HTTP status is 4xx/5xx # Parse the JSON response parsed_data <- jsonlite::fromJSON(httr::content(response, "text"), flatten = TRUE) # Convert NULL values to NA (tibbles don't allow NULL columns) tibble::tibble( ip = parsed_data$ip %||% NA_character_, country = parsed_data$country_name %||% NA_character_, region = parsed_data$region_name %||% NA_character_, city = parsed_data$city %||% NA_character_, zip = parsed_data$zip %||% NA_character_, latitude = parsed_data$latitude %||% NA_real_, longitude = parsed_data$longitude %||% NA_real_ ) }, error = function(e) { # Print error message for debugging, then return a row of NAs message(paste("Warning: Failed to process IP", ip_address, "-", e$message)) tibble::tibble( ip = ip_address, country = NA_character_, region = NA_character_, city = NA_character_, zip = NA_character_, latitude = NA_real_, longitude = NA_real_ ) }) }
Now run the batch processing with rate limiting (adjust Sys.sleep() based on ipstack's rate limits—free tier usually allows ~1 request/second):
# Replace with your actual IP vector and access key ip_addresses <- core_members$ip_address your_access_key <- "[YOUR_ACCESS_KEY_HERE]" # Process all IPs, adding a 1-second delay between requests ip_geodata <- ip_addresses %>% purrr::map_df(~{ Sys.sleep(1) ip_locate(.x, your_access_key) }) # View the results head(ip_geodata)
Key Notes for Your 1200 IPs
- Rate Limiting: ipstack enforces rate limits (even on paid tiers). Adding
Sys.sleep(1)prevents you from getting temporarily blocked, which is likely why you only got 8 results earlier—your requests were getting rejected after hitting the limit. - Error Handling: The
tryCatchensures that even if some IPs fail, the entire process doesn't crash, and you get a clear warning about which IPs had issues. - Simplification: Since you only need coordinates, you can strip down the
tibble()part of the function to justip,latitude, andlongitudeto keep things lean.
内容的提问来源于stack exchange,提问作者James R.

