You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中含嵌套df列表列的data.frame导出JSON时单行嵌套df转对象

嵌套tibble导出JSON去除多余数组包裹问题

问题场景

现有2个仅含1个列表列的tibble格式数据框,每个列表列存储1个单行的嵌套tibble:

  • 第一个数据框列名为longest_hw,嵌套tibble包含4个字段:
    • start:Date类型
    • end:Date类型
    • max_temp_in_hw:数值型
    • duration:数值型
      结构示例:
    # df1
    structure(list(longest_hw = list(structure(list(start = structure(12266, class = "Date"), 
        end = structure(12294, class = "Date"), max_temp_in_hw = 37.5, 
        duration = 28), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, 
    -1L)))), row.names = c(NA, -1L), class = c("tbl_df", "tbl", "data.frame"
    ))
    
    # 打印预览
    # A tibble: 1 × 1
    #   longest_hw      
    #   <list>          
    # 1 <tibble [1 × 4]>
    
  • 第二个数据框列名为highest_temp,嵌套tibble包含4个字段:
    • doy:整数型
    • cell_id:整数型
    • year:字符型
    • temp:数值型
      结构示例:
    structure(list(highest_temp = list(structure(list(doy = 220L, 
        cell_id = 68977L, year = "2013", temp = 39.3), class = c("tbl_df", 
    "tbl", "data.frame"), row.names = c(NA, -1L)))), row.names = c(NA, 
    -1L), class = c("tbl_df", "tbl", "data.frame"))
    
    # 打印预览
    # A tibble: 1 × 1
    #   highest_temp    
    #   <list>          
    # 1 <tibble [1 × 4]>
    

需求为合并两个顶层数据框后导出JSON文件:顶层列名longest_hw、highest_temp作为JSON一级键,对应值直接为JSON对象(以嵌套tibble的列名为二级键,对应字段值为值)。

原有代码问题

原有实现代码如下:

# 合并两个顶层data.frame
df = bind_cols(longest_hw, highest_temp)

# 转换为JSON格式
df_json = toJSON(df)

# 写入文件
write(df_json, "path_to_file")

导出结果中两个一级键的值为长度1的数组,对象被包裹在数组内,不符合预期:

[
  {
    "longest_hw": [
      {
        "start": "2003-08-02",
        "end": "2003-08-30",
        "max_temp_in_hw": 37.5,
        "duration": 28
      }
    ],
    "highest_temp": [
      { "doy": 220, "cell_id": 68977, "year": "2013", "temp": 39.3 }
    ]
  }
]

预期输出格式如下,一级键对应值直接为对象,无外层数组包裹:

[
  {
    "longest_hw": {
      "start": "2003-08-02",
      "end": "2003-08-30",
      "max_temp_in_hw": 37.5,
      "duration": 28
    },
    "highest_temp": {
      "doy": 220,
      "cell_id": 68977,
      "year": "2013",
      "temp": 39.3
    }
  }
]

解决方案

多余数组的来源有两个:一是列表列内存储的是tibble(data.frame类),jsonlite::toJSON默认将data.frame序列化为行对象组成的数组;二是默认规则下长度为1的向量会被序列化为长度1的数组。
处理步骤为:先将嵌套的单行tibble转为普通命名列表,再在序列化时开启单值自动拆箱参数即可。
可直接运行的代码如下:

library(dplyr)
library(jsonlite)

# 合并两个顶层表
df <- bind_cols(longest_hw, highest_temp)

# 将列表列内的单行嵌套tibble转换为普通命名列表
df <- df %>% 
  mutate(across(everything(), ~ lapply(.x, as.list)))

# 转换为JSON,自动拆箱单值,格式化输出,日期按ISO8601格式导出
df_json <- toJSON(
  df,
  auto_unbox = TRUE,
  pretty = TRUE,
  Date = "ISO8601"
)

# 写入文件
write(df_json, "path_to_output_file.json")

内容的提问来源于stack exchange,提问作者Lenn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 12:12:22