You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用rvest提取网页表格?Xpath方法遇问题求解决方案

解决Wunderground历史天气表格提取问题

1. 修复Xpath语法错误

你遇到的第一个错误是字符串引号冲突:R代码用双引号包裹Xpath表达式时,Xpath内部的双引号(如@id="inner")会导致R提前截断字符串。解决方式二选一:

  • 将包裹Xpath的引号改为单引号:
    webpage %>% html_element(xpath = '//*[@id="inner-content"]//table')
    
  • 转义Xpath内的双引号:
    webpage %>% html_element(xpath = "//*[@id=\"inner-content\"]//table")
    

2. 解决Full Xpath返回<NA>的问题

这个问题的核心是目标表格由JavaScript动态渲染,read_html()只能抓取页面静态初始代码,无法获取JS加载后的内容。以下是三种可行方案:

方案1:直接调用数据API(推荐)

通过浏览器开发者工具的「网络」面板,找到页面加载表格数据的JSON接口,直接请求数据:

library(jsonlite)
# 替换为对应日期的API地址(可从浏览器网络请求中复制)
api_url <- "https://api.weather.com/v1/location/KBNA:9:US/observations/historical.json?apiKey=e1f10a1e78da46f5b10a1e78da96f525&units=e&startDate=20240930&endDate=20240930"
weather_data <- fromJSON(api_url)
# 提取表格形式的数据
weather_table <- weather_data$observations

方案2:模拟浏览器渲染动态内容

使用RSelenium模拟浏览器加载页面,获取完整渲染后的HTML:

library(RSelenium)
library(rvest)
# 启动Chrome驱动(需提前安装ChromeDriver并配置环境变量)
driver <- rsDriver(browser = "chrome")
remDr <- driver[["client"]]
# 访问目标页面
remDr$navigate("https://www.wunderground.com/history/daily/KBNA/date/2024-9-30")
# 获取渲染后的页面源码
dynamic_webpage <- read_html(remDr$getPageSource()[[1]])
# 提取表格(可使用修正后的Xpath或直接解析表格)
weather_table <- dynamic_webpage %>% 
  html_element(xpath = '//*[@id="inner-content"]//table') %>% 
  html_table()
# 关闭浏览器驱动
remDr$close()
driver$server$stop()

方案3:尝试提取静态表格

若目标表格存在于静态HTML中,可跳过Xpath,直接用html_table()提取所有表格后筛选:

# 提取页面所有表格
all_tables <- webpage %>% html_table()
# 通常底部表格是最后一个,可通过查看结构确认
weather_table <- all_tables[[length(all_tables)]]

内容的提问来源于stack exchange,提问作者Vadim Katsemba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 19:15:11