添加完整请求头后网页仍返回403错误的爬取问题
解决requests请求目标站点返回403的问题
问题根源
你复制的请求头里的cookie包含Cloudflare反爬验证的核心字段(cf_clearance、__cf_bm),这些字段是浏览器通过人机验证后生成的,和你的当前浏览器会话、IP地址、User-Agent强绑定,且有较短的有效期,直接复制到requests中复用会触发Cloudflare的验证机制,导致返回403。另外if-modified-since字段会要求服务器返回未修改的内容,也可能干扰正常请求。
可行解决方案
方案1:使用cfscrape绕过Cloudflare验证
cfscrape是专门针对Cloudflare反爬的工具库,能自动处理验证流程:
- 安装依赖:
pip install cfscrape - 替换原有代码:
import cfscrape from bs4 import BeautifulSoup # 创建带Cloudflare处理的scraper scraper = cfscrape.create_scraper() link = 'https://m.happymh.com/manga/woshidashenxian' # 发送请求 response = scraper.get(link) print(response.status_code) # 解析页面 soup = BeautifulSoup(response.content, 'html.parser')
方案2:用Playwright模拟真实浏览器行为
模拟浏览器能完全复刻用户访问的流程,绕过大部分反爬机制,稳定性更高:
- 安装依赖:
pip install playwright playwright install chrome - 示例代码:
from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup with sync_playwright() as p: # 启动浏览器(headless=False可看到浏览器窗口,调试用;上线可改为True) browser = p.chromium.launch(headless=False) page = browser.new_page() # 访问目标页面 page.goto('https://m.happymh.com/manga/woshidashenxian') # 等待页面加载完成 page.wait_for_load_state('networkidle') # 获取页面源码 html_content = page.content() # 解析内容 soup = BeautifulSoup(html_content, 'html.parser') print(soup.title.text) # 关闭浏览器 browser.close()
方案3:精简请求头(辅助优化)
如果坚持用原生requests,先移除失效的cookie和限制字段,保留必要请求头尝试:
from bs4 import BeautifulSoup import requests headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.102 Safari/537.36', 'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'accept-language': 'en-US,en;q=0.9' } link = 'https://m.happymh.com/manga/woshidashenxian' req = requests.get(link, headers=headers) print(req.status_code)
注意:这种方法大概率还是会被Cloudflare拦截,仅作基础尝试。
内容的提问来源于stack exchange,提问作者Mohamad
相关产品推荐
相关产品推荐

