You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中将嵌套格式的TXT/JSON数据转换为规范的数据框

Hey there! Let’s figure out how to get this nested JSON data into R properly—your initial attempts with read.csv aren’t working because this isn’t actually a CSV file, it’s JSON Lines (one JSON object per line), which needs specialized tools to parse.

Why Your Initial Approaches Failed

  • Using read.csv() treats every character in the JSON as a CSV field, leading to the "more columns than column names" error.
  • Skipping row.names=NULL causes duplicate row name issues because R tries to use the messy first JSON column as row identifiers.

Step-by-Step Solution

Since you’re dealing with 2+ million rows, we’ll use tools optimized for large JSON datasets to avoid memory overload. Here are two reliable methods:

Method 1: Use ndjson (Fast for Large Files)

This package is built specifically for JSON Lines format, making it perfect for your dataset size:

# Install if you haven't already
install.packages("ndjson")
library(ndjson)

# Read the file (note: even if it's named .csv, it's JSON Lines)
raw_data <- ndjson::stream_in(file("test.csv"))

# Flatten nested columns (like the _id structure and the numeric type field)
flattened_data <- jsonlite::flatten(raw_data)

# Drop _id-related columns and keep your desired variables
final_data <- flattened_data[, !grepl("_id", names(flattened_data))]

# If you want to explicitly list the columns to keep (instead of excluding _id):
# desired_cols <- c("messageid", "attachments", "usernameid", "username", "server", "text", "type", "datetime", "type.$numberLong")
# final_data <- flattened_data[, desired_cols]

Method 2: Use jsonlite's Streaming Function

If you prefer jsonlite, its stream_in() function handles large files by processing rows incrementally:

install.packages("jsonlite")
library(jsonlite)

# Open a connection to your file
file_con <- file("test.csv")
open(file_con)

# Read and flatten the data in streams
flattened_data <- stream_in(file_con, flatten = TRUE)
close(file_con)

# Clean up to keep only your needed columns
final_data <- flattened_data[, !grepl("_id", names(flattened_data))]

Quick Notes on Edge Cases

  • Duplicate Keys: I noticed in your sample data, the second entry has duplicate usernameid keys (looks like a typo for username). JSON parsers will overwrite earlier keys with later ones, so you may want to fix that in the source file first if possible.
  • Memory: Both methods avoid loading the entire dataset into memory at once, which is critical for 2+ million rows.

内容的提问来源于stack exchange,提问作者Essan Rago

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 23:48:13