You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取网页失败(UseMethod错误)求助

问题原因分析

你遇到的Error in UseMethod("html_table"): no applicable method for 'html_table' applied to an object of class 'xml_missing'错误,本质是用html_node()/html_element()定位页面元素时未找到匹配节点,返回空的xml_missing对象,导致后续html_table()调用报错。

针对你爬取的MLB Savant页面,核心问题是:该页面表格数据通过JavaScript动态加载,直接用read_html()只能获取静态HTML源码,无法拿到动态渲染后的表格内容。


解决方案

思路1:直接调用数据接口(推荐)

页面导出CSV时的接口可直接返回结构化数据,绕过动态渲染问题:

library(tidyverse)
library(httr)

# 调用MLB官方数据接口
response <- GET("https://baseballsavant.mlb.com/leaderboard/percentile-rankings?type=batter&team=&csv=true")
hitting <- read_csv(content(response, as = "text"))

# 查看数据
head(hitting)

思路2:用无头浏览器渲染动态页面

如果必须解析动态渲染的HTML,可使用playwright模拟浏览器加载:

library(tidyverse)
library(rvest)
library(playwright)

# 初始化playwright环境
playwright_install()
pw <- playwright$launch()
browser <- pw$chromium$launch(headless = TRUE)
page <- browser$new_page()

# 加载页面并等待表格渲染完成
page$goto("https://baseballsavant.mlb.com/leaderboard/percentile-rankings?type=batter&team=")
page$wait_for_selector("table.table-savant")

# 提取并解析表格
html <- page$content()
hitting <- read_html(html) %>%
  html_node("table.table-savant") %>%
  html_table()

# 清理资源
browser$close()
pw$stop()

# 查看数据
head(hitting)

CSS选择器定位实用技巧

针对SelectorGadget效果有限的问题,给几个实用方法:

  • 直接在浏览器开发者工具(F12)的Elements面板,右键目标元素→Copy→Copy selector/Copy XPath,直接获取精准选择器
  • 优先用唯一标识的id或辨识度高的class(比如表格的table-savant类),避免嵌套过深的选择器
  • 用html_elements()(复数形式)替代html_node(),先查看所有匹配元素,确认是否存在目标节点

R网页爬取优质学习资源

  • 《R for Data Science》网页爬取章节:基础入门内容扎实
  • rvest官方文档:包含大量选择器用法与实战示例
  • 《Web Scraping with R》书籍:覆盖静态/动态页面爬取的各类场景
  • RStudio官方教程:有针对动态页面爬取的实战案例

内容的提问来源于stack exchange,提问作者Violin125

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 12:25:06