在R语言中解析JSON文件出错,附100条Yelp商家数据示例
Hey there! Let's fix that JSON parsing error you're facing with the Yelp business data in R. I’ve dealt with this exact dataset before, and the root cause is almost always those MongoDB-specific data types hiding in your JSON—things like ObjectId() and NumberInt() that standard JSON parsers don’t understand out of the box.
Common Issues in Your Yelp JSON
Looking at your sample data, here are the non-standard elements tripping up R:
ObjectId("5aab338ffc08b46adb7a2320"): MongoDB’s unique ID format, not valid standard JSONNumberInt(1): MongoDB’s integer type marker, which R doesn’t recognize as a numeric value&: HTML entity for&that might cause minor formatting issues with business names
Solution 1: Clean the JSON Text First
The simplest approach is to preprocess the raw JSON text to convert these non-standard types into valid JSON. Here’s how to do it step-by-step:
# Load required package library(jsonlite) # Read the raw JSON file as text lines yelp_raw_text <- readLines("your_yelp_business_data.json") # Replace MongoDB ObjectId with a plain string yelp_clean_text <- gsub('ObjectId\\("([^"]+)"\\)', '"\\1"', yelp_raw_text) # Replace NumberInt() with regular integers yelp_clean_text <- gsub('NumberInt\\((\\d+)\\)', '\\1', yelp_clean_text) # Optional: Fix HTML entities like & to & for readable business names yelp_clean_text <- gsub('&', '&', yelp_clean_text) # Now parse the cleaned JSON into an R data frame/list yelp_data <- fromJSON(paste(yelp_clean_text, collapse = "\n"))
Solution 2: Use a MongoDB-Specific Package
If you work with MongoDB data often, the mongolite package is built to handle these exact formats without manual cleaning:
# Install the package if you haven't already install.packages("mongolite") # Load the package and import the JSON directly library(mongolite) yelp_data <- mongoimport(file = "your_yelp_business_data.json")
Quick Tips to Avoid Headaches
- Test with a small sample first: Before processing all 100 entries, try the first 5-10 lines to make sure your cleaning/import works.
- Handle nested attributes: The
attributesfield is a nested list—usetidyr::unnest()if you want to flatten it into separate columns for easier analysis. - Check for other MongoDB types: If you run into
ISODate()or other markers, extend thegsub()logic to handle those similarly.
内容的提问来源于stack exchange,提问作者Nidhi Agarwal

