You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Indeed评论遇403错误,请求协助解决

Indeed评论爬取403问题解决方案

一、Indeed是否禁止爬取评论?

Indeed官方服务条款明确禁止未经许可的自动化爬取行为,尤其是针对用户评论这类内容的批量抓取。你收到的403状态码是平台反爬系统的拦截结果,说明你的请求被识别为非人工操作。

二、解决403错误的核心措施

  • 补充完整请求头:仅User-Agent不足以模拟真实浏览器,需添加Accept、Accept-Language、Referer等字段
  • 添加请求间隔:每次请求后暂停1-3秒,避免触发频率限制
  • 避免固定IP高频请求:必要时可使用代理IP轮换(注意合规性)
  • 优先考虑官方渠道:如果Indeed提供商家评论的API接口,使用API是最稳定合规的方式

三、修正后的爬取代码

from bs4 import BeautifulSoup
import pandas as pd
import requests
import numpy as np
import time

lst = []
# 模拟完整浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
    "Referer": "https://www.indeed.com/cmp/Meta-dd1502f2/reviews"
}

for i in range(0, 40, 20):
    print(f"处理页面偏移量: {i}")
    url = f'https://www.indeed.com/cmp/Meta-dd1502f2/reviews?start={i}'
    
    try:
        # 添加请求延迟
        time.sleep(2)
        page = requests.get(url, headers=headers)
        print(f'状态码: {page.status_code}')
        
        if page.status_code != 200:
            print(f"获取页面{i}失败,跳过")
            continue
            
        soup = BeautifulSoup(page.content, 'lxml')
        # 修正评论容器选择器(原选择器可能已失效)
        main_data = soup.find_all("div", class_="css-1qxtz39 eu4oa1w0")
        
        for data in main_data:
            date = np.nan
            try:
                author_info = data.find("span", attrs={"itemprop":"author"}).get_text(strip=True).split("-")
                if len(author_info) >=3:
                    date = author_info[2].strip()
            except AttributeError:
                pass
            
            title = np.nan
            try:
                title = data.find("h2", class_="css-1x93j7a e1tiznh50").get_text(strip=True)
            except AttributeError:
                pass
            
            status = np.nan
            location = np.nan
            try:
                author_info = data.find("span", attrs={"itemprop":"author"}).get_text(strip=True).split("-")
                if len(author_info)>=1:
                    status = author_info[0].strip()
                if len(author_info)>=2:
                    location = author_info[1].strip()
            except AttributeError:
                pass
            
            review = np.nan
            try:
                review = data.find("span", attrs={"itemprop":"reviewBody"}).get_text(strip=True)
            except AttributeError:
                pass
            
            pros = np.nan
            try:
                pros_section = data.find('h2', string="Pros")
                if pros_section:
                    pros = pros_section.next_sibling.get_text(strip=True)
            except:
                pass
            
            cons = np.nan
            try:
                cons_section = data.find('h2', string="Cons")
                if cons_section:
                    cons = cons_section.next_sibling.get_text(strip=True)
            except:
                pass
            
            rating = np.nan
            try:
                rating_elem = data.find("div", attrs={"itemprop":"reviewRating"}).find("button")
                if rating_elem:
                    rating = rating_elem['aria-label'].split(" ")[0]
            except AttributeError:
                pass
            
            lst.append([date, title, status, location, review, pros, cons, rating])
            
    except Exception as e:
        print(f"发生错误: {str(e)}")
        continue

df_meta = pd.DataFrame(data=lst, columns=['date', 'title', 'status', 'location', 'review', 'pros', 'cons', 'rating'])
print(df_meta.head())

四、代码说明

  1. 优化请求头,更贴近真实浏览器的请求特征
  2. 添加time.sleep(2)控制请求频率,降低被拦截风险
  3. 修正元素选择器,改用文本匹配(如"Pros"/"Cons")定位内容,避免因CSS类更新导致的失效
  4. 增加异常捕获与跳过机制,避免单个请求失败导致程序终止
  5. 对作者信息拆分增加长度判断,避免索引越界错误

内容的提问来源于stack exchange,提问作者Bharath Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 21:15:42