如何在R语言中将嵌套格式的TXT/JSON数据转换为规范的数据框
Hey there! Let’s figure out how to get this nested JSON data into R properly—your initial attempts with read.csv aren’t working because this isn’t actually a CSV file, it’s JSON Lines (one JSON object per line), which needs specialized tools to parse.
Why Your Initial Approaches Failed
- Using
read.csv()treats every character in the JSON as a CSV field, leading to the "more columns than column names" error. - Skipping
row.names=NULLcauses duplicate row name issues because R tries to use the messy first JSON column as row identifiers.
Step-by-Step Solution
Since you’re dealing with 2+ million rows, we’ll use tools optimized for large JSON datasets to avoid memory overload. Here are two reliable methods:
Method 1: Use ndjson (Fast for Large Files)
This package is built specifically for JSON Lines format, making it perfect for your dataset size:
# Install if you haven't already install.packages("ndjson") library(ndjson) # Read the file (note: even if it's named .csv, it's JSON Lines) raw_data <- ndjson::stream_in(file("test.csv")) # Flatten nested columns (like the _id structure and the numeric type field) flattened_data <- jsonlite::flatten(raw_data) # Drop _id-related columns and keep your desired variables final_data <- flattened_data[, !grepl("_id", names(flattened_data))] # If you want to explicitly list the columns to keep (instead of excluding _id): # desired_cols <- c("messageid", "attachments", "usernameid", "username", "server", "text", "type", "datetime", "type.$numberLong") # final_data <- flattened_data[, desired_cols]
Method 2: Use jsonlite's Streaming Function
If you prefer jsonlite, its stream_in() function handles large files by processing rows incrementally:
install.packages("jsonlite") library(jsonlite) # Open a connection to your file file_con <- file("test.csv") open(file_con) # Read and flatten the data in streams flattened_data <- stream_in(file_con, flatten = TRUE) close(file_con) # Clean up to keep only your needed columns final_data <- flattened_data[, !grepl("_id", names(flattened_data))]
Quick Notes on Edge Cases
- Duplicate Keys: I noticed in your sample data, the second entry has duplicate
usernameidkeys (looks like a typo forusername). JSON parsers will overwrite earlier keys with later ones, so you may want to fix that in the source file first if possible. - Memory: Both methods avoid loading the entire dataset into memory at once, which is critical for 2+ million rows.
内容的提问来源于stack exchange,提问作者Essan Rago
相关产品推荐
相关产品推荐

