Python爬取ZipRecruiter站点时出现HTTP 403 Forbidden错误如何解决?
问题原因
ZipRecruiter平台部署了反爬识别机制,默认使用requests发起的无标识请求会被判定为爬虫程序,返回的是人机验证页面而非正常的职位数据页面,因此你看到的内容看起来类似加密状态。
解决方案
按优先级依次尝试以下方案即可:
- 补充完整浏览器请求头,模拟真实用户访问
首先要在请求中添加浏览器的标识字段,优先补充User-Agent,可直接在自己的浏览器地址栏输入about:version查询当前浏览器的User-Agent值替换到示例中,示例代码如下:
from bs4 import BeautifulSoup import requests headers = { "User-Agent": "替换为你自己的浏览器User-Agent", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8" } response = requests.get( 'https://www.ziprecruiter.com/jobs-search?search=python&location=Ontario&company=&layout=&days=1', headers=headers ) # 状态码返回200代表请求基本被接受 print(response.status_code) soup = BeautifulSoup(response.text, 'html.parser') print(soup)
- 如果补充请求头后仍返回验证页面,可使用
cloudscraper库替代原生requests,该库会自动处理常见的Cloudflare等云防护验证逻辑,用法和requests基本一致:
先执行安装命令:pip install cloudscraper
示例代码如下:
from bs4 import BeautifulSoup import cloudscraper scraper = cloudscraper.create_scraper() response = scraper.get('https://www.ziprecruiter.com/jobs-search?search=python&location=Ontario&company=&layout=&days=1') soup = BeautifulSoup(response.text, 'html.parser') print(soup)
- 若以上方案都不生效,可使用无头浏览器模拟真人操作,比如
undetected-chromedriver,完全还原浏览器的行为特征,可绕过绝大多数反爬检测。
额外提醒:爬取时注意控制请求频率,每次请求间隔至少2秒,避免短时间发起大量请求导致IP被封禁。
内容的提问来源于stack exchange,提问作者Henul
相关产品推荐
相关产品推荐

