如何用rvest提取EDI网站Andrews LTER数据集的Package Id?
解决方法
首先检查目标页面的HTML结构会发现,Package Id所在单元格的class是id而非Package Id,且每个Package Id都包裹在<a>标签中,调整选择器即可正确提取:
# Load required libraries library(rvest) library(dplyr) # Define the URL of the website url <- "http://portal.edirepository.org:80/nis/simpleSearch?defType=edismax&q=*:*&fq=-scope:ecotrends&fq=-scope:lter-landsat*&fq=scope:(knb-lter-and)&fl=id,packageid,title,author,organization,pubdate,coordinates&debug=false" # Read the HTML content from the website page <- read_html(url) # Extract Package Ids correctly packageIds <- page %>% html_nodes("td.id a") %>% # 选择class为id的td下的a标签 html_text() # 验证提取结果的数量(应为147个) length(packageIds)
也可以直接选择td.id提取文本,效果一致:
packageIds <- page %>% html_nodes("td.id") %>% html_text()
若需要确认提取内容,可打印前几个元素:
head(packageIds)
内容的提问来源于stack exchange,提问作者tassones
相关产品推荐
相关产品推荐

