You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用R语言结合rvest库从指定HTML页面提取Zip格式URL

解决R语言rvest提取Zip格式URL的问题

嘿,我看到你在使用rvest提取S3页面里的Zip链接时遇到了错误,这是因为你用错了函数——html_attrs()是用来获取元素的全部属性的,它不需要传入属性名参数,所以你传"href"就会触发那个"unused argument"错误。

下面是正确的实现步骤和代码,帮你顺利提取所有目标Zip链接:

正确代码示例

# 加载所需库
library(rvest)

# 目标页面链接
base_link <- "https://divvy-tripdata.s3.amazonaws.com/index.html"

# 读取页面内容
page_html <- read_html(base_link)

# 1. 选中所有<a>标签元素
# 2. 提取每个<a>标签的href属性
# 3. 过滤出以.zip结尾的链接
zip_links <- page_html %>%
  html_elements("a") %>%
  html_attr("href") %>%
  grep("\\.zip$", ., value = TRUE)

# 如果页面里的href是相对路径,拼接成完整可访问的URL
full_zip_links <- paste0("https://divvy-tripdata.s3.amazonaws.com/", zip_links)

# 查看结果
print(full_zip_links)

代码解释

运行这段代码后,你就能得到页面里所有的Zip格式链接了。

内容的提问来源于stack exchange,提问作者Stephen Aung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 16:22:28