You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R获取UNDP网页中PDF与DOCX文件的下载链接?

解决UNDP评估页面文件下载问题

这类页面的文件下载通常依赖后端API接口而非可见的<a>标签,你可以通过以下步骤获取真实下载链接并批量下载:

步骤1:分析真实下载接口

打开目标页面(比如https://erc.undp.org/evaluation/evaluations/detail/7834?tab=documents),按F12打开浏览器开发者工具:

  • 切换到「Network」标签
  • 点击页面上的下载按钮,观察触发的网络请求,找到类似/evaluation/api/documents/{文档ID}/download的请求URL,这就是真实的下载接口。

步骤2:提取页面中的文档ID

页面的HTML里通常会隐藏文档ID(比如通过自定义属性存储),用rvest提取:

library(rvest)
library(dplyr)

# 目标页面URL
target_url <- "https://erc.undp.org/evaluation/evaluations/detail/7834?tab=documents"
page_html <- read_html(target_url)

# 提取所有文档ID(根据页面实际属性调整,比如data-document-id)
doc_ids <- page_html %>%
  html_nodes("[data-document-id]") %>%
  html_attr("data-document-id")

# 过滤空值
doc_ids <- doc_ids[!is.na(doc_ids)]

步骤3:构造下载链接并批量下载

结合API接口格式生成完整下载URL,再用httr保持会话(这类网站需要Cookie验证)下载:

library(httr)

# 初始化会话,获取页面Cookie
session <- html_session(target_url)

# API下载接口前缀(根据抓包结果调整)
base_download_api <- "https://erc.undp.org/evaluation/api/documents/%s/download"

# 遍历每个文档ID,生成链接并下载
for (id in doc_ids) {
  download_url <- sprintf(base_download_api, id)
  # 发送请求
  resp <- GET(session, download_url)
  
  # 从响应头提取文件名
  disp_header <- resp$headers$`content-disposition`
  filename <- sub('attachment; filename="(.*)"', '\\1', disp_header)
  
  # 保存文件
  writeBin(content(resp, "raw"), filename)
  cat(sprintf("已下载:%s\n", filename))
}

注意事项

  • 如果抓包发现下载是POST请求,需要调整代码用POST()并携带对应的表单参数(比如文档ID)
  • 若页面需要登录,需先通过html_session()模拟登录流程,再执行下载操作

内容的提问来源于stack exchange,提问作者Arihant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:45:20