You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言rvest爬取Glassdoor评论子评分并存入数据框问题求解

问题背景

使用rvest爬取Glassdoor网站评论时,需要将每条评论整理为data.frame的一行,其中工作生活平衡、文化与价值观等5项子评分需要作为单独列存储。原代码中子评分提取段无法运行,且未处理单条评论部分子评分为空的场景。

原代码核心错误
  • 语法错误:R中序列写法错误,for(i in 1 to 5)是非法语法,正确写法为1:5;最终拼接data.frame时culture_values与后续列之间遗漏逗号
  • 逻辑错误:循环中i是数值类型,不是HTML节点对象,直接对数值调用html_element必然触发报错;子评分赋值写在评论循环内部、数据框拼接写在循环外部,最终只会保留最后一条评论的子评分结果
  • 鲁棒性不足:未做子评分缺失处理,强行按固定索引取子评分,遇到缺项时会出现值错位、长度不匹配报错;CSS类与评分的映射逻辑未做空值判断,容易匹配到无效值
  • 选择器错误:子评分模块的DOM节点定位不准,没有区分评分标签和评分星条的节点
可直接运行的修正代码
## 加载依赖包
library(httr)  
library(xml2)  
library(rvest) 
library(purrr) 
library(tidyverse)
library(lubridate)

## 目标页面地址
url = "https://www.glassdoor.com/Reviews/Google-Reviews-E9079.htm"
pg_reviews = read_html(url)

## 预生成CSS类与评分值的映射表
class.ratings = c()
styles = pg_reviews %>% html_elements('style')
for(s in styles) {
  class_attr = s %>% html_attr('data-emotion-css')
  if(!is.na(class_attr)){
    class_name = paste0('css-', class_attr)
    rating_match = str_match(s %>% html_text2(), '(\\d+)%')[2]
    if(!is.na(rating_match)){
      class.ratings[class_name] = as.numeric(rating_match)/20
    }
  }
}

## 提取所有评论节点
reviews = pg_reviews %>% html_elements('.gdReview')

## 逐评论解析字段,直接生成结构化数据框
Google_reviews = map_dfr(reviews, function(re){
  # 提取基础字段
  review_summary = re %>% html_element(".reviewLink") %>% html_text2()
  overall_rating = re %>% html_element(".mr-xsm") %>% html_text2()
  review_pros = re %>% html_element(".v2__EIReviewDetailsV2__fullWidth:nth-child(1) span") %>% html_text2()
  review_cons = re %>% html_element(".v2__EIReviewDetailsV2__fullWidth:nth-child(2) span") %>% html_text2()
  
  # 初始化5个维度子评分的默认值为NA
  wlb = NA_real_
  culture = NA_real_
  career = NA_real_
  comp = NA_real_
  mgmt = NA_real_
  
  # 提取当前评论下存在的所有子评分
  subrating_items = re %>% html_elements('.subRatings__SubRating__container li')
  if(length(subrating_items) > 0){
    walk(subrating_items, function(item){
      # 提取子评分维度名称
      item_label = item %>% html_element('div:first-child') %>% html_text2()
      # 匹配对应评分值
      rating_class = item %>% html_element('.ratingBar') %>% html_attr('class') %>% str_split(' ') %>% pluck(1,1)
      item_score = class.ratings[rating_class] %>% unname()
      
      # 按维度名称赋值到对应列
      case_when(
        str_detect(item_label, "Work Life Balance") ~ wlb <<- item_score,
        str_detect(item_label, "Culture & Values") ~ culture <<- item_score,
        str_detect(item_label, "Career Opportunities") ~ career <<- item_score,
        str_detect(item_label, "Compensation and Benefits") ~ comp <<- item_score,
        str_detect(item_label, "Senior Management") ~ mgmt <<- item_score
      )
    })
  }
  
  # 返回单条评论的结构化行数据
  tibble(
    summary = review_summary,
    overall_rating = overall_rating,
    pros = review_pros,
    cons = review_cons,
    work_life_balance = wlb,
    culture_values = culture,
    career_opportunities = career,
    comp_benefits = comp,
    senior_management = mgmt
  )
})
修正逻辑说明
  • 放弃全局按位置提取字段的方案,改为逐评论节点解析所有字段,从根源上避免不同评论字段数不一致导致的错位
  • 子评分提取不依赖固定位置:先解析当前评论下所有存在的子评分,按标签文本匹配到对应维度,缺失的维度自动赋值为NA,完全适配部分子评分空缺的情况
  • 用purrr::map_dfr替代手动for循环加rbind的写法,自动按列名拼接所有评论结果,不会遗漏数据
  • 给CSS评分映射、节点提取环节都加了空值判断,遇到页面结构微小变动时不会直接报错终止
  • 补充说明:如果直接调用read_html拿到的页面数据不全,是Glassdoor反爬机制拦截导致,需要用httr::GET添加合法的User-Agent请求头后再解析页面内容

内容的提问来源于stack exchange,提问作者HDobbins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 05:45:40