Python爬取Indeed评论遇403错误,请求协助解决
Indeed评论爬取403问题解决方案
一、Indeed是否禁止爬取评论?
Indeed官方服务条款明确禁止未经许可的自动化爬取行为,尤其是针对用户评论这类内容的批量抓取。你收到的403状态码是平台反爬系统的拦截结果,说明你的请求被识别为非人工操作。
二、解决403错误的核心措施
- 补充完整请求头:仅
User-Agent不足以模拟真实浏览器,需添加Accept、Accept-Language、Referer等字段 - 添加请求间隔:每次请求后暂停1-3秒,避免触发频率限制
- 避免固定IP高频请求:必要时可使用代理IP轮换(注意合规性)
- 优先考虑官方渠道:如果Indeed提供商家评论的API接口,使用API是最稳定合规的方式
三、修正后的爬取代码
from bs4 import BeautifulSoup import pandas as pd import requests import numpy as np import time lst = [] # 模拟完整浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Referer": "https://www.indeed.com/cmp/Meta-dd1502f2/reviews" } for i in range(0, 40, 20): print(f"处理页面偏移量: {i}") url = f'https://www.indeed.com/cmp/Meta-dd1502f2/reviews?start={i}' try: # 添加请求延迟 time.sleep(2) page = requests.get(url, headers=headers) print(f'状态码: {page.status_code}') if page.status_code != 200: print(f"获取页面{i}失败,跳过") continue soup = BeautifulSoup(page.content, 'lxml') # 修正评论容器选择器(原选择器可能已失效) main_data = soup.find_all("div", class_="css-1qxtz39 eu4oa1w0") for data in main_data: date = np.nan try: author_info = data.find("span", attrs={"itemprop":"author"}).get_text(strip=True).split("-") if len(author_info) >=3: date = author_info[2].strip() except AttributeError: pass title = np.nan try: title = data.find("h2", class_="css-1x93j7a e1tiznh50").get_text(strip=True) except AttributeError: pass status = np.nan location = np.nan try: author_info = data.find("span", attrs={"itemprop":"author"}).get_text(strip=True).split("-") if len(author_info)>=1: status = author_info[0].strip() if len(author_info)>=2: location = author_info[1].strip() except AttributeError: pass review = np.nan try: review = data.find("span", attrs={"itemprop":"reviewBody"}).get_text(strip=True) except AttributeError: pass pros = np.nan try: pros_section = data.find('h2', string="Pros") if pros_section: pros = pros_section.next_sibling.get_text(strip=True) except: pass cons = np.nan try: cons_section = data.find('h2', string="Cons") if cons_section: cons = cons_section.next_sibling.get_text(strip=True) except: pass rating = np.nan try: rating_elem = data.find("div", attrs={"itemprop":"reviewRating"}).find("button") if rating_elem: rating = rating_elem['aria-label'].split(" ")[0] except AttributeError: pass lst.append([date, title, status, location, review, pros, cons, rating]) except Exception as e: print(f"发生错误: {str(e)}") continue df_meta = pd.DataFrame(data=lst, columns=['date', 'title', 'status', 'location', 'review', 'pros', 'cons', 'rating']) print(df_meta.head())
四、代码说明
- 优化请求头,更贴近真实浏览器的请求特征
- 添加
time.sleep(2)控制请求频率,降低被拦截风险 - 修正元素选择器,改用文本匹配(如"Pros"/"Cons")定位内容,避免因CSS类更新导致的失效
- 增加异常捕获与跳过机制,避免单个请求失败导致程序终止
- 对作者信息拆分增加长度判断,避免索引越界错误
内容的提问来源于stack exchange,提问作者Bharath Kumar
相关产品推荐
相关产品推荐

