You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取今日美国MLB薪资数据表格?R语言rvest报错求助

解决MLB 2022薪资数据抓取问题

原代码报错是因为目标页面未使用标准<table>标签,而是用div元素模拟表格结构。以下是第一页数据的抓取方案:

代码实现

library(rvest)
library(dplyr)

# 读取目标页面
page <- read_html("https://databases.usatoday.com/mlb-salaries-2022/")

# 提取表格表头
headers <- page %>% 
  html_nodes(".datatable-header-cell") %>% 
  html_text() %>% 
  trimws()

# 提取每一行的单元格数据
rows <- page %>% 
  html_nodes(".datatable-row") %>% 
  lapply(function(row_node) {
    row_node %>% 
      html_nodes(".datatable-cell") %>% 
      html_text() %>% 
      trimws()
  })

# 转换为数据框并设置列名
salary_data <- do.call(rbind.data.frame, rows)
colnames(salary_data) <- headers

# 查看数据前几行
head(salary_data)

关键说明

  • 页面的表头对应类选择器.datatable-header-cell,数据行对应.datatable-row,每个单元格对应.datatable-cell
  • 使用trimws()清理文本前后的空白字符,避免数据混乱
  • 若后续遇到动态加载的分页数据,可结合RSelenium模拟浏览器渲染获取内容

内容的提问来源于stack exchange,提问作者user2006134

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 10:25:23