You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言rvest搭配SelectorGadget抓取财富500强员工数报错排查

初始代码返回空值的原因

你最开始写的CSS选择器抓取方案拿不到数据,核心原因有两个:

  • 财富网站的公司详情页是JavaScript动态渲染的,rvest::read_html()只能抓取页面初始返回的静态HTML源码,你在浏览器里用SelectorGadget看到的员工数字段,是浏览器加载完页面后执行JS脚本才渲染出来的,静态源码里根本不存在对应class的节点,选择器自然匹配不到内容,返回character(0)。
  • 你用的选择器里的class后缀(比如7f9lE、2AHH7)是前端构建自动生成的随机哈希值,会随网站版本更新随时变化,就算用支持JS渲染的爬虫工具,硬写这个选择器也很容易失效。
更新后代码报错的修复方法

你遇到的length(idx) > 0 is not TRUE报错,和stopifnot()的校验逻辑没关系,不要删这段校验——这段代码的作用是找不到对应节点时抛出明确错误,避免后续取子元素时出现更难排查的空值错误。
报错的根源是你写死了从预加载JSON里找数据的嵌套路径:不同公司的页面结构存在细微差异,不是所有页面都严格按照body→company-about-wrapper→company-information的层级存数据,硬编码路径很容易断链。
最稳妥的方案是写个递归函数遍历整个预加载JSON,直接定位employees字段,不用管具体嵌套层级,适配所有页面的结构差异,同时加错误捕获和请求延时,避免被反爬拦截、单页失败导致整个循环中断。

修复后的完整可运行代码如下:

# 先加载依赖包,没装的先跑install.packages(c("rvest", "dplyr", "jsonlite"))
library(rvest)
library(dplyr)
library(jsonlite)

# 待抓取的财富500强公司链接
company_urls <- c("https://fortune.com/company/walmart/", "https://fortune.com/company/amazon-com/"              
,"https://fortune.com/company/apple/"                   
,"https://fortune.com/company/cvs-health/"              
,"https://fortune.com/company/unitedhealth-group/"      
, "https://fortune.com/company/berkshire-hathaway/"      
, "https://fortune.com/company/mckesson/"                
,"https://fortune.com/company/amerisourcebergen/"       
, "https://fortune.com/company/alphabet/"                
, "https://fortune.com/company/exxon-mobil/"             
,"https://fortune.com/company/att/"                     
,"https://fortune.com/company/costco/"                  
,"https://fortune.com/company/cigna/"                   
, "https://fortune.com/company/cardinal-health/"         
,"https://fortune.com/company/microsoft/"               
,"https://fortune.com/company/walgreens-boots-alliance/"
,"https://fortune.com/company/kroger/"                  
, "https://fortune.com/company/home-depot/"              
,"https://fortune.com/company/jpmorgan-chase/"          
,"https://fortune.com/company/verizon/"                 
,"https://fortune.com/company/ford-motor/"              
, "https://fortune.com/company/general-motors/"          
,"https://fortune.com/company/anthem/"                  
, "https://fortune.com/company/centene/"                 
,"https://fortune.com/company/fannie-mae/"              
, "https://fortune.com/company/comcast/"                 
, "https://fortune.com/company/chevron/"                 
,"https://fortune.com/company/dell-technologies/"       
,"https://fortune.com/company/bank-of-america-corp/"    
,"https://fortune.com/company/target/")

# 递归遍历列表查找员工数字段的辅助函数
find_employee_field <- function(list_obj) {
  if (is.list(list_obj)) {
    # 当前层级存在目标字段直接返回
    if (!is.null(list_obj$employees)) return(list_obj$employees)
    # 不存在就递归遍历所有子节点
    for (child in list_obj) {
      res <- find_employee_field(child)
      if (!is.null(res)) return(res)
    }
  }
  # 遍历完找不到就返回NA
  return(NA_character_)
}

# 初始化结果存储向量
employee_counts <- character(length(company_urls))

for (i in seq_along(company_urls)){
  tryCatch({
    # 提取页面预加载的JSON数据
    preload_json <- read_html(
      company_urls[i],
      # 加浏览器UA降低被反爬拦截概率
      user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
    ) |>
      html_element("script#preload") |> 
      html_text() |>
      sub("\\s*window\\.__PRELOADED_STATE__ = ", "", x = _, perl = TRUE) |>
      sub(";\\s*$", "", x = _, perl = TRUE) |>
      fromJSON(simplifyVector = FALSE)
    
    # 递归查找员工数字段
    employee_counts[i] <- find_employee_field(preload_json)
    # 加1秒请求延时,避免请求太频繁被封
    Sys.sleep(1)
    message(sprintf("抓取完成:%s,员工数:%s", company_urls[i], employee_counts[i]))
  }, error = function(e) {
    employee_counts[i] <<- NA_character_
    message(sprintf("抓取失败:%s,原因:%s", company_urls[i], e$message))
  })
}

# 整理成数据框
final_result <- tibble(
  url = company_urls,
  employee_count = employee_counts
)

补充说明

  • 抓回来的employee_count是带千分位分隔符的字符串,后续做量化分析时,自己去掉逗号转成数值类型即可。
  • 如果遇到个别页面抓取失败,多半是临时的网络波动或者反爬拦截,单独重跑对应链接即可,不用修改核心逻辑。

内容的提问来源于stack exchange,提问作者Xian Zhao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 03:30:50