You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的rvest包从NCBI页面提取指定S3链接?

使用rvest从NCBI Trace页面表格提取S3链接

要提取目标S3链接,你需要定位到页面中的表格元素,再提取对应行的链接内容。以下是修正后的代码:

library(rvest)

# 目标URL
url <- "https://trace.ncbi.nlm.nih.gov/Traces/index.html?view=run_browser&acc=SRR11393390&display=data-access"

# 读取页面内容
html_content <- read_html(url)

# 定位表格并提取S3链接
s3_links <- html_content %>%
  # 定位页面内的数据表格
  html_nodes("#ph-maincontent table") %>%
  # 提取表格中的所有<a>标签
  html_nodes("a") %>%
  # 获取链接的href属性值
  html_attr("href") %>%
  # 筛选出以s3://开头的目标链接
  grep("^s3://", ., value = TRUE)

# 输出结果
print(s3_links)

关键说明:

  • 页面内的S3链接都包裹在表格的<a>标签中,通过html_nodes("a")定位这些元素。
  • html_attr("href")用于提取标签中的实际链接地址。
  • 用grep过滤出s3://前缀的链接,精准匹配你需要的资源地址。

如果页面表格有更具体的ID或类名(比如class="data-table"),可以将html_nodes("#ph-maincontent table")替换为更精准的选择器,提升定位效率。

内容的提问来源于stack exchange,提问作者A4747

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:05:17