You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取ZipRecruiter站点时出现HTTP 403 Forbidden错误如何解决?

问题原因

ZipRecruiter平台部署了反爬识别机制,默认使用requests发起的无标识请求会被判定为爬虫程序,返回的是人机验证页面而非正常的职位数据页面,因此你看到的内容看起来类似加密状态。

解决方案

按优先级依次尝试以下方案即可:

  • 补充完整浏览器请求头,模拟真实用户访问
    首先要在请求中添加浏览器的标识字段,优先补充User-Agent,可直接在自己的浏览器地址栏输入about:version查询当前浏览器的User-Agent值替换到示例中,示例代码如下:
from bs4 import BeautifulSoup
import requests

headers = {
    "User-Agent": "替换为你自己的浏览器User-Agent",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8"
}

response = requests.get(
    'https://www.ziprecruiter.com/jobs-search?search=python&location=Ontario&company=&layout=&days=1',
    headers=headers
)
# 状态码返回200代表请求基本被接受
print(response.status_code)
soup = BeautifulSoup(response.text, 'html.parser')
print(soup)
  • 如果补充请求头后仍返回验证页面,可使用cloudscraper库替代原生requests,该库会自动处理常见的Cloudflare等云防护验证逻辑,用法和requests基本一致:
    先执行安装命令:pip install cloudscraper
    示例代码如下:
from bs4 import BeautifulSoup
import cloudscraper

scraper = cloudscraper.create_scraper()
response = scraper.get('https://www.ziprecruiter.com/jobs-search?search=python&location=Ontario&company=&layout=&days=1')
soup = BeautifulSoup(response.text, 'html.parser')
print(soup)
  • 若以上方案都不生效,可使用无头浏览器模拟真人操作,比如undetected-chromedriver,完全还原浏览器的行为特征,可绕过绝大多数反爬检测。

额外提醒:爬取时注意控制请求频率,每次请求间隔至少2秒,避免短时间发起大量请求导致IP被封禁。

内容的提问来源于stack exchange,提问作者Henul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:45:02