如何用rvest提取网页表格?Xpath方法遇问题求解决方案
解决Wunderground历史天气表格提取问题
1. 修复Xpath语法错误
你遇到的第一个错误是字符串引号冲突:R代码用双引号包裹Xpath表达式时,Xpath内部的双引号(如@id="inner")会导致R提前截断字符串。解决方式二选一:
- 将包裹Xpath的引号改为单引号:
webpage %>% html_element(xpath = '//*[@id="inner-content"]//table') - 转义Xpath内的双引号:
webpage %>% html_element(xpath = "//*[@id=\"inner-content\"]//table")
2. 解决Full Xpath返回<NA>的问题
这个问题的核心是目标表格由JavaScript动态渲染,read_html()只能抓取页面静态初始代码,无法获取JS加载后的内容。以下是三种可行方案:
方案1:直接调用数据API(推荐)
通过浏览器开发者工具的「网络」面板,找到页面加载表格数据的JSON接口,直接请求数据:
library(jsonlite) # 替换为对应日期的API地址(可从浏览器网络请求中复制) api_url <- "https://api.weather.com/v1/location/KBNA:9:US/observations/historical.json?apiKey=e1f10a1e78da46f5b10a1e78da96f525&units=e&startDate=20240930&endDate=20240930" weather_data <- fromJSON(api_url) # 提取表格形式的数据 weather_table <- weather_data$observations
方案2:模拟浏览器渲染动态内容
使用RSelenium模拟浏览器加载页面,获取完整渲染后的HTML:
library(RSelenium) library(rvest) # 启动Chrome驱动(需提前安装ChromeDriver并配置环境变量) driver <- rsDriver(browser = "chrome") remDr <- driver[["client"]] # 访问目标页面 remDr$navigate("https://www.wunderground.com/history/daily/KBNA/date/2024-9-30") # 获取渲染后的页面源码 dynamic_webpage <- read_html(remDr$getPageSource()[[1]]) # 提取表格(可使用修正后的Xpath或直接解析表格) weather_table <- dynamic_webpage %>% html_element(xpath = '//*[@id="inner-content"]//table') %>% html_table() # 关闭浏览器驱动 remDr$close() driver$server$stop()
方案3:尝试提取静态表格
若目标表格存在于静态HTML中,可跳过Xpath,直接用html_table()提取所有表格后筛选:
# 提取页面所有表格 all_tables <- webpage %>% html_table() # 通常底部表格是最后一个,可通过查看结构确认 weather_table <- all_tables[[length(all_tables)]]
内容的提问来源于stack exchange,提问作者Vadim Katsemba
相关产品推荐
相关产品推荐

