You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中提取复杂嵌套JSON文件中的assetID与filename元素

解决方案

核心原因

你当前获取的metadata_published是嵌套的JSON格式字符串,仅完成了外层JSON的解析,需要对该字符串做二次JSON解析,才能提取内部存储的附件assetID、文件名等字段。

完整实现代码

# 加载依赖包
library(RJSONIO)
library(dplyr)
library(openxlsx)

# 读取外层元数据JSON
df_raw <- RJSONIO::fromJSON("https://healthdata.gov/resource/6hii-ae4f.json", flatten=TRUE)
# 提取内层嵌套的JSON字符串
metadata_json_str <- df_raw[[1]][["metadata_published"]]
# 二次解析内层JSON字符串为可操作的列表结构
metadata_parsed <- RJSONIO::fromJSON(metadata_json_str)

# 提取所有附件的assetID与文件名,转为数据框
asset_df <- lapply(metadata_parsed$attachments, function(item) {
  data.frame(
    asset_id = item$assetId,
    file_name = item$name,
    stringsAsFactors = FALSE
  )
}) %>% bind_rows()

# 筛选符合命名规则的目标xlsx报告文件
target_file_df <- asset_df %>% 
  filter(grepl("^Community_Profile_Report_\\d{8}_Public\\.xlsx$", file_name))

# ----------------------
# 后续批量读取处理逻辑
# ----------------------
# 下载链接模板
download_url_tpl <- "https://healthdata.gov/api/views/gqxm-d9w9/files/%s?download=true&filename=%s"

# 批量迭代处理所有目标文件(此处用lapply符合你提到的apply类函数需求)
process_result <- lapply(1:nrow(target_file_df), function(i) {
  # 拼接当前文件的合法下载链接,对文件名做URL编码避免特殊字符报错
  current_url <- sprintf(download_url_tpl, 
                         target_file_df$asset_id[i], 
                         URLencode(target_file_df$file_name[i]))
  # 读取并按你的规则处理数据
  temp_df <- read.xlsx(current_url, sheet = 6) %>%
    select(1, 2, 6, 8, 16, 76, 77) %>%
    slice(-1)
  colnames(temp_df) <- c("county", "FIPS", "state", "pop", "cases_last7", "vaccinated", "vacc_prop")
  num_cols <- c("FIPS", "pop", "cases_last7", "vaccinated", "vacc_prop")
  temp_df[num_cols] <- sapply(temp_df[num_cols], as.numeric)
  temp_df <- temp_df %>%
    filter(complete.cases(.)) %>%
    mutate(cases_last7_100k = cases_last7 / pop * 100000,
           vacc_prop = vacc_prop * 100) %>%
    filter(state != "PR" & vacc_prop < 100)
  return(temp_df)
})

# 给结果列表命名为对应的文件名
names(process_result) <- target_file_df$file_name

注意事项

  • 拼接下载链接时必须使用URLencode()处理文件名,否则文件名中的空格、特殊字符会导致请求失败
  • 批量处理时可添加tryCatch()逻辑捕获个别文件的读取异常,避免单文件失败中断全量处理流程
  • 最终target_file_df即为你需要的assetID和文件名对应数据框,process_result为所有文件处理后的结果列表

内容的提问来源于stack exchange,提问作者David Crow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 16:24:03