使用R语言rvest爬取Glassdoor评论子评分并存入数据框问题求解
问题背景
使用rvest爬取Glassdoor网站评论时,需要将每条评论整理为data.frame的一行,其中工作生活平衡、文化与价值观等5项子评分需要作为单独列存储。原代码中子评分提取段无法运行,且未处理单条评论部分子评分为空的场景。
原代码核心错误
- 语法错误:R中序列写法错误,
for(i in 1 to 5)是非法语法,正确写法为1:5;最终拼接data.frame时culture_values与后续列之间遗漏逗号 - 逻辑错误:循环中
i是数值类型,不是HTML节点对象,直接对数值调用html_element必然触发报错;子评分赋值写在评论循环内部、数据框拼接写在循环外部,最终只会保留最后一条评论的子评分结果 - 鲁棒性不足:未做子评分缺失处理,强行按固定索引取子评分,遇到缺项时会出现值错位、长度不匹配报错;CSS类与评分的映射逻辑未做空值判断,容易匹配到无效值
- 选择器错误:子评分模块的DOM节点定位不准,没有区分评分标签和评分星条的节点
可直接运行的修正代码
## 加载依赖包 library(httr) library(xml2) library(rvest) library(purrr) library(tidyverse) library(lubridate) ## 目标页面地址 url = "https://www.glassdoor.com/Reviews/Google-Reviews-E9079.htm" pg_reviews = read_html(url) ## 预生成CSS类与评分值的映射表 class.ratings = c() styles = pg_reviews %>% html_elements('style') for(s in styles) { class_attr = s %>% html_attr('data-emotion-css') if(!is.na(class_attr)){ class_name = paste0('css-', class_attr) rating_match = str_match(s %>% html_text2(), '(\\d+)%')[2] if(!is.na(rating_match)){ class.ratings[class_name] = as.numeric(rating_match)/20 } } } ## 提取所有评论节点 reviews = pg_reviews %>% html_elements('.gdReview') ## 逐评论解析字段,直接生成结构化数据框 Google_reviews = map_dfr(reviews, function(re){ # 提取基础字段 review_summary = re %>% html_element(".reviewLink") %>% html_text2() overall_rating = re %>% html_element(".mr-xsm") %>% html_text2() review_pros = re %>% html_element(".v2__EIReviewDetailsV2__fullWidth:nth-child(1) span") %>% html_text2() review_cons = re %>% html_element(".v2__EIReviewDetailsV2__fullWidth:nth-child(2) span") %>% html_text2() # 初始化5个维度子评分的默认值为NA wlb = NA_real_ culture = NA_real_ career = NA_real_ comp = NA_real_ mgmt = NA_real_ # 提取当前评论下存在的所有子评分 subrating_items = re %>% html_elements('.subRatings__SubRating__container li') if(length(subrating_items) > 0){ walk(subrating_items, function(item){ # 提取子评分维度名称 item_label = item %>% html_element('div:first-child') %>% html_text2() # 匹配对应评分值 rating_class = item %>% html_element('.ratingBar') %>% html_attr('class') %>% str_split(' ') %>% pluck(1,1) item_score = class.ratings[rating_class] %>% unname() # 按维度名称赋值到对应列 case_when( str_detect(item_label, "Work Life Balance") ~ wlb <<- item_score, str_detect(item_label, "Culture & Values") ~ culture <<- item_score, str_detect(item_label, "Career Opportunities") ~ career <<- item_score, str_detect(item_label, "Compensation and Benefits") ~ comp <<- item_score, str_detect(item_label, "Senior Management") ~ mgmt <<- item_score ) }) } # 返回单条评论的结构化行数据 tibble( summary = review_summary, overall_rating = overall_rating, pros = review_pros, cons = review_cons, work_life_balance = wlb, culture_values = culture, career_opportunities = career, comp_benefits = comp, senior_management = mgmt ) })
修正逻辑说明
- 放弃全局按位置提取字段的方案,改为逐评论节点解析所有字段,从根源上避免不同评论字段数不一致导致的错位
- 子评分提取不依赖固定位置:先解析当前评论下所有存在的子评分,按标签文本匹配到对应维度,缺失的维度自动赋值为
NA,完全适配部分子评分空缺的情况 - 用
purrr::map_dfr替代手动for循环加rbind的写法,自动按列名拼接所有评论结果,不会遗漏数据 - 给CSS评分映射、节点提取环节都加了空值判断,遇到页面结构微小变动时不会直接报错终止
- 补充说明:如果直接调用
read_html拿到的页面数据不全,是Glassdoor反爬机制拦截导致,需要用httr::GET添加合法的User-Agent请求头后再解析页面内容
内容的提问来源于stack exchange,提问作者HDobbins
相关产品推荐
相关产品推荐

