You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python requests爬取网页遇Mod_Security Not Acceptable报错解决

报错现象

爬取目标站点时未返回预期页面内容,服务器返回406状态拦截响应,响应内容如下:

b'<head><title>Not Acceptable!</title></head><body><h1>Not Acceptable!</h1><p>An appropriate representation of the requested resource could not be found on this server. This error was generated by Mod_Security.</p></body></html>'

原始代码运行输出参考:
代码运行报错输出

报错根因

该拦截由目标站点部署的Mod_Security WAF防火墙触发:requests库默认发起请求时,请求头内的User-Agent字段会携带python-requests相关标识,站点反爬规则识别到该类非浏览器的爬虫特征后,会直接拦截请求,拒绝返回正常页面内容。

修复方案

为请求添加真实浏览器的请求头,模拟普通用户浏览器的访问特征,即可绕过这层基础反爬拦截。
修复后的代码:

from bs4 import BeautifulSoup
import requests
import pandas as pd

url = 'https://insights.blackcoffer.com/how-is-login-logout-time-tracking-for-employees-in-office-done-by-ai/'

# 模拟Chrome桌面端浏览器的常规请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8"
}

resp = requests.get(url, headers=headers)
# 验证请求状态,返回200即为请求成功
print(resp.status_code)
page = resp.content
注意事项
  • 若添加基础请求头后仍被拦截,可适当增加请求间隔,避免短时间内高频发起请求触发更严格的拦截规则
  • 爬取行为需遵守目标站点的robots协议,请勿将爬取内容用于违规商业用途
  • 若后续遇到会话校验类拦截,可使用requests.Session()对象维持访问会话,模拟真实用户的访问路径

内容的提问来源于stack exchange,提问作者Ashwin Pillai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 13:03:30