如何使用rvest爬取目标网页的Tracing Information表格数据?
获取Tracing Information表格数据的解决方案
可以通过XPath精准定位到Tracing Information标题对应的表格,替代之前按索引取表格的方式,这样更稳定且不易受页面结构变化影响。以下是修改后的代码:
library(rvest) url <- "https://www.aaacooper.com/pwb/Transit/ProTrackResults.aspx?ProNum=241939875&AllAccounts=true" page <- read_html(url) # 定位Tracing Information表格并解析 tracing_info_table <- page %>% # 通过XPath找到标题后的表格 html_nodes(xpath = "//*[text()='Tracing Information']/following-sibling::table") %>% # 解析为表格,将第一行设为表头 html_table(header = TRUE) %>% # 取第一个(也是唯一的)表格 .[[1]] # 查看结果 View(tracing_info_table)
代码说明
- 使用XPath
//*[text()='Tracing Information']/following-sibling::table直接定位到标题紧跟的表格,避免了按索引取表格(如.[[2]])可能失效的问题; header = TRUE参数确保表格的第一行被识别为列名,最终输出的表格结构更符合需求。
内容的提问来源于stack exchange,提问作者Jaskeil
相关产品推荐
相关产品推荐

