You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用rvest或同类R包自动下载UNICEF指标数据?

解决UNICEF指标数据自动下载问题

一、先解决指标代码识别问题

你原代码的问题是管道符缺失(%>应改为%>%),且仅提取了页面节点但未解析其中的指标链接与代码。以下代码可批量抓取所有指标的名称、链接及对应的指标代码:

library(rvest)
library(dplyr)

# 指标列表页面地址
page_url <- "https://data.unicef.org/indicator-profile/"
page <- read_html(page_url)

# 提取所有指标的核心信息
indicators <- page %>%
  html_nodes("div.data-inner-wrapper a") %>%
  tibble(
    indicator_name = html_text(.),
    indicator_link = html_attr("href"),
    indicator_code = basename(indicator_link) # 从链接末尾提取指标代码
  )

# 查看前5条结果验证
head(indicators)

运行后会得到包含所有指标代码的数据集,比如你提到的CME_TMY5T9会出现在indicator_code列中。

二、单个指标数据下载

基于指标代码,可编写函数自动进入对应指标页面,抓取下载链接并完成数据下载:

# 定义单个指标下载函数
download_unicef_indicator <- function(indicator_code, save_path = "./") {
  # 构造指标页面地址
  indicator_page_url <- paste0("https://data.unicef.org/indicator/", indicator_code, "/")
  indicator_page <- read_html(indicator_page_url)
  
  # 优先抓取CSV格式下载链接
  download_link <- indicator_page %>%
    html_nodes(xpath = "//a[contains(@class, 'download-button') and contains(@href, '.csv')]") %>%
    html_attr("href") %>%
    first()
  
  # 如果没有CSV,尝试抓取Excel格式
  if (is.na(download_link)) {
    download_link <- indicator_page %>%
      html_nodes(xpath = "//a[contains(@class, 'download-button') and contains(@href, '.xlsx')]") %>%
      html_attr("href") %>%
      first()
  }
  
  # 执行下载
  if (!is.na(download_link)) {
    # 根据文件格式设置保存后缀
    file_ext <- ifelse(grepl(".csv", download_link), ".csv", ".xlsx")
    dest_file <- paste0(save_path, indicator_code, file_ext)
    
    download.file(
      url = download_link,
      destfile = dest_file,
      mode = "wb"
    )
    message(paste0("✅ 成功下载指标 ", indicator_code, " 的数据"))
  } else {
    message(paste0("❌ 未找到指标 ", indicator_code, " 的下载链接"))
  }
}

# 测试下载5-9岁死亡人数数据
download_unicef_indicator("CME_TMY5T9")

三、按国家批量下载全量数据

针对指定国家(如阿富汗AFG),可批量抓取该国所有指标的数据:

# 定义国家批量下载函数
download_unicef_country_data <- function(country_code, save_path = "./") {
  # 创建保存目录(如果不存在)
  if (!dir.exists(save_path)) dir.create(save_path)
  
  # 构造国家页面地址
  country_page_url <- paste0("https://data.unicef.org/country/", country_code, "/")
  country_page <- read_html(country_page_url)
  
  # 提取该国所有指标的代码
  country_indicators <- country_page %>%
    html_nodes("div.data-inner-wrapper a") %>%
    tibble(indicator_link = html_attr("href")) %>%
    mutate(indicator_code = basename(indicator_link)) %>%
    distinct(indicator_code) # 去重避免重复下载
  
  # 遍历下载每个指标
  for (code in country_indicators$indicator_code) {
    download_unicef_indicator(code, save_path)
    Sys.sleep(1) # 添加1秒延迟,避免请求频率过高被限制
  }
}

# 测试下载阿富汗(AFG)的所有指标数据
download_unicef_country_data("AFG", "./afghanistan_data/")

注意事项

  • 部分指标可能没有公开的下载数据,函数会自动提示未找到链接
  • 不要频繁大量请求,建议保留Sys.sleep()延迟,避免被网站封禁IP
  • 如果遇到反爬限制,可尝试添加请求头(使用httr包的add_headers())模拟浏览器请求

内容的提问来源于stack exchange,提问作者thehand0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 03:55:31