如何用rvest或同类R包自动下载UNICEF指标数据?
解决UNICEF指标数据自动下载问题
一、先解决指标代码识别问题
你原代码的问题是管道符缺失(%>应改为%>%),且仅提取了页面节点但未解析其中的指标链接与代码。以下代码可批量抓取所有指标的名称、链接及对应的指标代码:
library(rvest) library(dplyr) # 指标列表页面地址 page_url <- "https://data.unicef.org/indicator-profile/" page <- read_html(page_url) # 提取所有指标的核心信息 indicators <- page %>% html_nodes("div.data-inner-wrapper a") %>% tibble( indicator_name = html_text(.), indicator_link = html_attr("href"), indicator_code = basename(indicator_link) # 从链接末尾提取指标代码 ) # 查看前5条结果验证 head(indicators)
运行后会得到包含所有指标代码的数据集,比如你提到的CME_TMY5T9会出现在indicator_code列中。
二、单个指标数据下载
基于指标代码,可编写函数自动进入对应指标页面,抓取下载链接并完成数据下载:
# 定义单个指标下载函数 download_unicef_indicator <- function(indicator_code, save_path = "./") { # 构造指标页面地址 indicator_page_url <- paste0("https://data.unicef.org/indicator/", indicator_code, "/") indicator_page <- read_html(indicator_page_url) # 优先抓取CSV格式下载链接 download_link <- indicator_page %>% html_nodes(xpath = "//a[contains(@class, 'download-button') and contains(@href, '.csv')]") %>% html_attr("href") %>% first() # 如果没有CSV,尝试抓取Excel格式 if (is.na(download_link)) { download_link <- indicator_page %>% html_nodes(xpath = "//a[contains(@class, 'download-button') and contains(@href, '.xlsx')]") %>% html_attr("href") %>% first() } # 执行下载 if (!is.na(download_link)) { # 根据文件格式设置保存后缀 file_ext <- ifelse(grepl(".csv", download_link), ".csv", ".xlsx") dest_file <- paste0(save_path, indicator_code, file_ext) download.file( url = download_link, destfile = dest_file, mode = "wb" ) message(paste0("✅ 成功下载指标 ", indicator_code, " 的数据")) } else { message(paste0("❌ 未找到指标 ", indicator_code, " 的下载链接")) } } # 测试下载5-9岁死亡人数数据 download_unicef_indicator("CME_TMY5T9")
三、按国家批量下载全量数据
针对指定国家(如阿富汗AFG),可批量抓取该国所有指标的数据:
# 定义国家批量下载函数 download_unicef_country_data <- function(country_code, save_path = "./") { # 创建保存目录(如果不存在) if (!dir.exists(save_path)) dir.create(save_path) # 构造国家页面地址 country_page_url <- paste0("https://data.unicef.org/country/", country_code, "/") country_page <- read_html(country_page_url) # 提取该国所有指标的代码 country_indicators <- country_page %>% html_nodes("div.data-inner-wrapper a") %>% tibble(indicator_link = html_attr("href")) %>% mutate(indicator_code = basename(indicator_link)) %>% distinct(indicator_code) # 去重避免重复下载 # 遍历下载每个指标 for (code in country_indicators$indicator_code) { download_unicef_indicator(code, save_path) Sys.sleep(1) # 添加1秒延迟,避免请求频率过高被限制 } } # 测试下载阿富汗(AFG)的所有指标数据 download_unicef_country_data("AFG", "./afghanistan_data/")
注意事项
- 部分指标可能没有公开的下载数据,函数会自动提示未找到链接
- 不要频繁大量请求,建议保留
Sys.sleep()延迟,避免被网站封禁IP - 如果遇到反爬限制,可尝试添加请求头(使用
httr包的add_headers())模拟浏览器请求
内容的提问来源于stack exchange,提问作者thehand0
相关产品推荐
相关产品推荐

