You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法通过bs4抓取Glassdoor评论分项评分,需识别星星颜色判定评分

Glassdoor评论分项评分抓取方案

Glassdoor的星星评分不会通过元素结构区分亮灭,而是靠CSS类或内联样式标记。不用识别颜色,直接抓这些标记或者隐藏的数值即可:

方法1:通过CSS类判断点亮的星星

先找到评论里的分项评分区域,再遍历星星元素,统计带点亮标记的数量:

import requests
from bs4 import BeautifulSoup

url = 'https://www.glassdoor.com/Reviews/Walmart-Reviews-E715_P2.htm?filter.iso3Language=eng'
# 加请求头避免被拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位单条评论的分项评分(需根据实际页面结构调整选择器)
review = soup.find(class_="gdReview")
if review:
    rating_items = review.find_all(class_="review-details__item")
    for item in rating_items:
        label = item.find(class_="review-details__label").text.strip()
        # 替换为实际星星的类名,这里假设点亮的星星带css-1q6a3jt类
        filled_stars = len(item.find_all(class_="css-1q6a3jt"))
        print(f"{label}: {filled_stars}/5")

方法2:直接抓取隐藏的评分数值

Glassdoor通常会把评分存在data-rating这类属性里,比判断星星更可靠:

# 抓取整体评分
overall_rating = soup.find(class_="ratingNumber mr-xsm")['data-rating']
print(f"整体评分: {overall_rating}/5")

# 抓取分项评分
for item in soup.find_all(class_="review-details__item"):
    # 找到带data-rating的评分元素
    rating_val = item.find(class_="ratingNumber")['data-rating']
    label = item.find(class_="review-details__label").text.strip()
    print(f"{label}: {rating_val}/5")

关键提醒

  • Glassdoor反爬严格,纯requests容易被封,建议搭配selenium模拟浏览器,或者使用代理、更新请求头。
  • 页面类名会不定期更新,要随时检查源码调整选择器。

内容的提问来源于stack exchange,提问作者Jaevapple

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 00:10:48