You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取网页表格时变量为空的问题求助

问题原因与解决方法

问题根源

你代码里的CSS选择器.td:nth-child(1)、.td:nth-child(2)等是错误的——网页中的<td>标签并没有td这个class属性,所以rvest找不到任何匹配的节点,自然返回空向量,最终生成的data frame就是0观测值。

修复后的代码

直接使用html_table()函数可以更简单准确地提取表格,无需手动逐个抓取列:

library(tidyverse)
library(rvest)

# 读取目标网页
url <- read_html("https://www.zyxware.com/articles/5363/list-of-fortune-500-companies-and-their-websites-2015")

# 定位目标表格并转换为data frame
fortune500 <- url %>%
  # 通过summary属性精准定位表格(避免抓取页面其他无关表格)
  html_node(xpath = "//table[contains(@summary, 'Fortune 500 Companies List')]") %>%
  # 自动识别表头并转换表格
  html_table(header = TRUE)

# 查看结果
head(fortune500)

补充说明

如果一定要用手动抓取列的方式,需要修正CSS选择器,同时注意剔除表头行:

# 先定位到目标表格
target_table <- url %>% html_node(xpath = "//table[contains(@summary, 'Fortune 500 Companies List')]")

# 抓取各列
rank <- target_table %>% html_nodes("tr td:nth-child(1)") %>% html_text()
company <- target_table %>% html_nodes("tr td:nth-child(2)") %>% html_text()
website <- target_table %>% html_nodes("tr td:nth-child(3)") %>% html_text()

# 去掉表头行(html_text会包含表头文本,需剔除)
fortune500 <- data.frame(
  rank = rank[-1],
  company = company[-1],
  website = website[-1]
)

内容的提问来源于stack exchange,提问作者Taylor Foster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 17:13:12