R语言rvest搭配SelectorGadget抓取财富500强员工数报错排查
初始代码返回空值的原因
你最开始写的CSS选择器抓取方案拿不到数据,核心原因有两个:
- 财富网站的公司详情页是JavaScript动态渲染的,
rvest::read_html()只能抓取页面初始返回的静态HTML源码,你在浏览器里用SelectorGadget看到的员工数字段,是浏览器加载完页面后执行JS脚本才渲染出来的,静态源码里根本不存在对应class的节点,选择器自然匹配不到内容,返回character(0)。 - 你用的选择器里的class后缀(比如
7f9lE、2AHH7)是前端构建自动生成的随机哈希值,会随网站版本更新随时变化,就算用支持JS渲染的爬虫工具,硬写这个选择器也很容易失效。
更新后代码报错的修复方法
你遇到的length(idx) > 0 is not TRUE报错,和stopifnot()的校验逻辑没关系,不要删这段校验——这段代码的作用是找不到对应节点时抛出明确错误,避免后续取子元素时出现更难排查的空值错误。
报错的根源是你写死了从预加载JSON里找数据的嵌套路径:不同公司的页面结构存在细微差异,不是所有页面都严格按照body→company-about-wrapper→company-information的层级存数据,硬编码路径很容易断链。
最稳妥的方案是写个递归函数遍历整个预加载JSON,直接定位employees字段,不用管具体嵌套层级,适配所有页面的结构差异,同时加错误捕获和请求延时,避免被反爬拦截、单页失败导致整个循环中断。
修复后的完整可运行代码如下:
# 先加载依赖包,没装的先跑install.packages(c("rvest", "dplyr", "jsonlite")) library(rvest) library(dplyr) library(jsonlite) # 待抓取的财富500强公司链接 company_urls <- c("https://fortune.com/company/walmart/", "https://fortune.com/company/amazon-com/" ,"https://fortune.com/company/apple/" ,"https://fortune.com/company/cvs-health/" ,"https://fortune.com/company/unitedhealth-group/" , "https://fortune.com/company/berkshire-hathaway/" , "https://fortune.com/company/mckesson/" ,"https://fortune.com/company/amerisourcebergen/" , "https://fortune.com/company/alphabet/" , "https://fortune.com/company/exxon-mobil/" ,"https://fortune.com/company/att/" ,"https://fortune.com/company/costco/" ,"https://fortune.com/company/cigna/" , "https://fortune.com/company/cardinal-health/" ,"https://fortune.com/company/microsoft/" ,"https://fortune.com/company/walgreens-boots-alliance/" ,"https://fortune.com/company/kroger/" , "https://fortune.com/company/home-depot/" ,"https://fortune.com/company/jpmorgan-chase/" ,"https://fortune.com/company/verizon/" ,"https://fortune.com/company/ford-motor/" , "https://fortune.com/company/general-motors/" ,"https://fortune.com/company/anthem/" , "https://fortune.com/company/centene/" ,"https://fortune.com/company/fannie-mae/" , "https://fortune.com/company/comcast/" , "https://fortune.com/company/chevron/" ,"https://fortune.com/company/dell-technologies/" ,"https://fortune.com/company/bank-of-america-corp/" ,"https://fortune.com/company/target/") # 递归遍历列表查找员工数字段的辅助函数 find_employee_field <- function(list_obj) { if (is.list(list_obj)) { # 当前层级存在目标字段直接返回 if (!is.null(list_obj$employees)) return(list_obj$employees) # 不存在就递归遍历所有子节点 for (child in list_obj) { res <- find_employee_field(child) if (!is.null(res)) return(res) } } # 遍历完找不到就返回NA return(NA_character_) } # 初始化结果存储向量 employee_counts <- character(length(company_urls)) for (i in seq_along(company_urls)){ tryCatch({ # 提取页面预加载的JSON数据 preload_json <- read_html( company_urls[i], # 加浏览器UA降低被反爬拦截概率 user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" ) |> html_element("script#preload") |> html_text() |> sub("\\s*window\\.__PRELOADED_STATE__ = ", "", x = _, perl = TRUE) |> sub(";\\s*$", "", x = _, perl = TRUE) |> fromJSON(simplifyVector = FALSE) # 递归查找员工数字段 employee_counts[i] <- find_employee_field(preload_json) # 加1秒请求延时,避免请求太频繁被封 Sys.sleep(1) message(sprintf("抓取完成:%s,员工数:%s", company_urls[i], employee_counts[i])) }, error = function(e) { employee_counts[i] <<- NA_character_ message(sprintf("抓取失败:%s,原因:%s", company_urls[i], e$message)) }) } # 整理成数据框 final_result <- tibble( url = company_urls, employee_count = employee_counts )
补充说明
- 抓回来的
employee_count是带千分位分隔符的字符串,后续做量化分析时,自己去掉逗号转成数值类型即可。 - 如果遇到个别页面抓取失败,多半是临时的网络波动或者反爬拦截,单独重跑对应链接即可,不用修改核心逻辑。
内容的提问来源于stack exchange,提问作者Xian Zhao
相关产品推荐
相关产品推荐

