You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中解析JSON文件出错,附100条Yelp商家数据示例

Hey there! Let's fix that JSON parsing error you're facing with the Yelp business data in R. I’ve dealt with this exact dataset before, and the root cause is almost always those MongoDB-specific data types hiding in your JSON—things like ObjectId() and NumberInt() that standard JSON parsers don’t understand out of the box.

Common Issues in Your Yelp JSON

Looking at your sample data, here are the non-standard elements tripping up R:

  • ObjectId("5aab338ffc08b46adb7a2320"): MongoDB’s unique ID format, not valid standard JSON
  • NumberInt(1): MongoDB’s integer type marker, which R doesn’t recognize as a numeric value
  • &: HTML entity for & that might cause minor formatting issues with business names

Solution 1: Clean the JSON Text First

The simplest approach is to preprocess the raw JSON text to convert these non-standard types into valid JSON. Here’s how to do it step-by-step:

# Load required package
library(jsonlite)

# Read the raw JSON file as text lines
yelp_raw_text <- readLines("your_yelp_business_data.json")

# Replace MongoDB ObjectId with a plain string
yelp_clean_text <- gsub('ObjectId\\("([^"]+)"\\)', '"\\1"', yelp_raw_text)

# Replace NumberInt() with regular integers
yelp_clean_text <- gsub('NumberInt\\((\\d+)\\)', '\\1', yelp_clean_text)

# Optional: Fix HTML entities like &amp; to & for readable business names
yelp_clean_text <- gsub('&amp;', '&', yelp_clean_text)

# Now parse the cleaned JSON into an R data frame/list
yelp_data <- fromJSON(paste(yelp_clean_text, collapse = "\n"))

Solution 2: Use a MongoDB-Specific Package

If you work with MongoDB data often, the mongolite package is built to handle these exact formats without manual cleaning:

# Install the package if you haven't already
install.packages("mongolite")

# Load the package and import the JSON directly
library(mongolite)
yelp_data <- mongoimport(file = "your_yelp_business_data.json")

Quick Tips to Avoid Headaches

  • Test with a small sample first: Before processing all 100 entries, try the first 5-10 lines to make sure your cleaning/import works.
  • Handle nested attributes: The attributes field is a nested list—use tidyr::unnest() if you want to flatten it into separate columns for easier analysis.
  • Check for other MongoDB types: If you run into ISODate() or other markers, extend the gsub() logic to handle those similarly.

内容的提问来源于stack exchange,提问作者Nidhi Agarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:53:08