You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言展开数据框中多层JSON列的技术求助

解析数据框中多层JSON字符串并展开为多列的解决方案

问题背景

数据框的response列存储着一段多层JSON数组字符串,需要将其解析并展开为数据框的多列。以下是该JSON字符串示例:

"[
{\"question\":
 {\"id\":314123,
  \"text\":\"What did you see in this module?\",
  \"question_id\":151,
  \"requirements\":
  {\"answer_type\":\"text\",
   \"text_required\":true}
  },
 \"response\":
 {\"free_text\":
  {\"id\":\"314123_ft\",
   \"text\":\"Identifying critical knowledge\"}
 }
},
{\"question\":
 {\"id\":314124,
  \"text\":\"What do you feel is the gap?\",
  \"question_id\":152,
  \"requirements\":
  {\"answer_type\":\"text\",
   \"text_required\":true}
 },
 \"response\":
  {\"free_text\":
   {\"id\":\"314124_ft\",
    \"text\":\"Knowledge is clearly mapped out\"}
 }
},
{\"question\":
 {\"id\":314125,
  \"text\":\"Which of the modules you need to focus on\",
  \"question_id\":153,
  \"requirements\":
  {\"answer_type\":\"text\",
   \"text_required\":true}
 },
 \"response\":
 {\"free_text\":
  {\"id\":\"3141125_ft\",
   \"text\":\"preparation of key knowledge\"}
 }
}
]"

尝试过以下两种方法,均报错:

方法一及报错

fromJSON(df) %>% 
  unnest(c(response))

报错信息:

Error in (function (classes, fdef, mtable)  : 
  unable to find an inherited method for function ‘fromJSON’ for signature ‘"data.frame", "missing"’

方法二及报错

df <- df %>% 
  rowwise() %>%
  do(data.frame(fromJSON(.$response, flatten = T))) %>%
  ungroup() %>%
  bind_cols(df %>% select(-response))

报错信息:

Error in (function (..., row.names = NULL, check.rows = FALSE, check.names = TRUE,  : 
  arguments imply differing number of rows: 1, 0

解决方案

核心问题分析

  • fromJSON函数需要传入单个字符串,而非整个数据框,直接传入df会导致类型不匹配。
  • 你的JSON是数组结构,解析后返回列表,需要逐行处理并正确展开嵌套层级,再与原数据合并。

正确代码实现

使用jsonlite+dplyr+tidyr组合处理:

library(jsonlite)
library(dplyr)
library(tidyr)
library(stringr) # 用于清理JSON字符串(如果需要)

# 步骤1:清理JSON字符串(如果存在首尾多余双引号)
df <- df %>%
  mutate(
    response = str_remove_all(response, '^"|"$')
  )

# 步骤2:解析每行的JSON字符串为扁平列表
df <- df %>%
  mutate(response_parsed = map(response, ~ fromJSON(.x, flatten = TRUE)))

# 步骤3:逐层展开嵌套列表为数据框列
df_expanded <- df %>%
  unnest_wider(response_parsed) %>%
  unnest_wider(question) %>%
  unnest_wider(response) %>%
  unnest_wider(free_text) %>%
  # 移除原始response列(可选)
  select(-response)

代码说明

  1. JSON清理:如果JSON字符串首尾带有多余的双引号,用str_remove_all去掉,避免解析失败。
  2. 逐行解析:map(response, ~ fromJSON(.x, flatten = TRUE))对每行的response字符串单独解析,flatten=TRUE自动将嵌套的JSON结构转为扁平列表。
  3. 层级展开:unnest_wider()逐层展开列表列,把每个嵌套层级的字段转为数据框的独立列,依次处理response_parsed、question、response、free_text。

验证结果

执行后,df_expanded会包含以下字段:

  • 原数据框除response外的所有列
  • 解析后的id、text、question_id、requirements.answer_type、requirements.text_required
  • free_text.id、free_text.text(扁平化后会直接显示为free_text.id这类列名)

内容的提问来源于stack exchange,提问作者Unai Vicente

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 07:07:07