将含不同嵌套数据的JSON文件转换为统一DataFrame
处理JSON嵌套数据:提取指定字段并统一格式
需求
从目标JSON数据中,保留所有用户共有的userId、username、fullName字段,同时提取latestPosts中存在的locationName和locationId;对于无latestPosts或该字段下无位置信息的条目,缺失值用NA填充,生成结构统一的DataFrame。
完整代码
library(jsonlite) library(dplyr) library(purrr) # 读取数据 test_file <- fromJSON("https://raw.githubusercontent.com/datacfb123/testdata/main/test_file_lp.json") # 提取并整理字段 result_df <- test_file %>% select(userId, username, fullName) %>% mutate( # 提取每个用户latestPosts中的位置字段,无对应字段时返回空 loc_info = map(latestPosts, ~ select(.x, any_of(c("locationName", "locationId")))), # 提取locationName,无数据则填NA locationName = map_chr(loc_info, ~ if (nrow(.x) > 0) .x$locationName else NA_character_), # 提取locationId,无数据则填NA locationId = map_chr(loc_info, ~ if (nrow(.x) > 0) .x$locationId else NA_character_) ) %>% # 移除临时列和原始嵌套字段 select(-loc_info, -latestPosts) # 输出结果 print(result_df)
代码解释
any_of():避免因某些latestPosts无位置字段而报错,自动忽略不存在的字段map()+map_chr():遍历每个用户的嵌套数据,提取目标值,缺失时填充NA- 最终只保留需要的字段,确保结果结构统一
输出示例
生成的DataFrame结构如下:
| userId | username | fullName | locationName | locationId |
|---|---|---|---|---|
| 1 | username1 | User One | NA | NA |
| 2 | username2 | User Two | NA | NA |
| 3 | username3 | User Three | New York | 12345 |
内容的提问来源于stack exchange,提问作者wizkids121
相关产品推荐
相关产品推荐

