如何使用rvest提取表格每行的href属性及所有链接?
使用rvest提取HTML表格中的链接
没问题,用rvest处理这个需求非常直接,我给你两种常用的实现方式,按需选择:
步骤1:准备工作
首先确保你已经安装并加载了rvest包,如果还没安装,先执行这一步:
install.packages("rvest") library(rvest)
步骤2:解析HTML内容
把你提供的HTML内容定义为字符串(如果是本地文件或在线网页,也可以直接用read_html("文件路径/网址")读取):
html_content <- ' <!DOCTYPE html> <html> <body> <table> <tr> <td> <a href="https://www.r-project.org/">R</a><br> <a href="https://www.rstudio.com/">RStudio</a> </td> </tr> <tr> <td> <a href="https://community.rstudio.com/">Rstudio Community</a> </td> </tr> </table> </body> </html> ' # 解析HTML page <- read_html(html_content)
方式1:提取所有链接(不区分行)
如果只需要把表格里的所有链接一次性提取出来,用这个方法:
# 选中所有<a>标签,提取href属性 all_links <- page %>% html_elements("a") %>% html_attr("href") # 查看结果 all_links
输出结果:
[1] "https://www.r-project.org/" "https://www.rstudio.com/" [3] "https://community.rstudio.com/"
方式2:按表格行分组提取链接
如果需要保留“每行对应哪些链接”的结构,就先选中所有表格行(<tr>标签),再逐行提取链接:
# 需要用到purrr包(rvest通常会和它搭配使用) library(purrr) # 遍历每一行,提取该行内的所有链接 links_by_row <- page %>% html_elements("tr") %>% map(function(row) { row %>% html_elements("a") %>% html_attr("href") }) # 查看结果 links_by_row
输出结果:
[[1]] [1] "https://www.r-project.org/" "https://www.rstudio.com/" [[2]] [1] "https://community.rstudio.com/"
简单解释下核心函数:
html_elements():用来定位HTML中的目标标签(比如"a"选中所有链接,"tr"选中所有表格行)html_attr("href"):提取标签的指定属性值,这里就是链接地址map():用来遍历每一行,批量处理每行的链接提取
内容的提问来源于stack exchange,提问作者rjss
相关产品推荐
相关产品推荐

