You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从httr请求返回的Header列表中提取多组status字段?

解决方法

方法1:直接提取status子字段构建单行数据框

直接指定要提取的status子字段,和url一起组成单行tibble,确保map_dfr可以正确合并:

pacman::p_load(httr, dplyr, purrr, tibble)

some_urls <- c("https://www.psychologytoday.com/us/therapists/new-york/a?page=10",
               "https://www.psychologytoday.com/us/therapists/new-york/a?page=4",
               "https://www.psychologytoday.com/us/therapists/new-york/a?page=140",
               "https://www.psychologytoday.com/us/therapists/new-york/a?page=3"
)

df <- map_dfr(some_urls, ~{
  resp <- httr::GET(.x)
  tibble(
    url = resp$url,
    status_category = resp$status$category,
    status_reason = resp$status$reason,
    status_message = resp$status$message
  )
})

方法2:先嵌套存储status再展开

如果需要灵活处理status的所有子字段,可以先将status作为嵌套列表列存储,再用unnest_wider展开:

df <- map_dfr(some_urls, ~{
  resp <- httr::GET(.x)
  tibble(
    url = resp$url,
    status = list(resp$status)
  )
}) %>%
  tidyr::unnest_wider(status)

方法3:结合magrittr管道调整提取逻辑

如果坚持用管道风格,可将status转换为tibble后和url列绑定:

df <- map_dfr(some_urls, ~{
  httr::GET(.x) %>%
    {tibble(url = .$url) %>% dplyr::bind_cols(tibble::as_tibble(.$status))}
})

错误原因

报错的核心是httr响应对象的status是嵌套命名列表(包含category、reason等子元素),而url是单个字符串。直接用magrittr::extract(c("url", "status"))返回的结构是list(url="xxx", status=list(...)),map_dfr自动合并时会错误地将url重复匹配到status的每个子元素行,最终导致合并后出现多个重复的url列,触发唯一列名校验错误。

内容的提问来源于stack exchange,提问作者Anna Jones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 20:57:17