You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加完整请求头后网页仍返回403错误的爬取问题

解决requests请求目标站点返回403的问题

问题根源

你复制的请求头里的cookie包含Cloudflare反爬验证的核心字段(cf_clearance、__cf_bm),这些字段是浏览器通过人机验证后生成的,和你的当前浏览器会话、IP地址、User-Agent强绑定,且有较短的有效期,直接复制到requests中复用会触发Cloudflare的验证机制,导致返回403。另外if-modified-since字段会要求服务器返回未修改的内容,也可能干扰正常请求。

可行解决方案

方案1:使用cfscrape绕过Cloudflare验证

cfscrape是专门针对Cloudflare反爬的工具库,能自动处理验证流程:

  1. 安装依赖:
    pip install cfscrape
    
  2. 替换原有代码:
    import cfscrape
    from bs4 import BeautifulSoup
    
    # 创建带Cloudflare处理的scraper
    scraper = cfscrape.create_scraper()
    link = 'https://m.happymh.com/manga/woshidashenxian'
    # 发送请求
    response = scraper.get(link)
    print(response.status_code)
    # 解析页面
    soup = BeautifulSoup(response.content, 'html.parser')
    

方案2:用Playwright模拟真实浏览器行为

模拟浏览器能完全复刻用户访问的流程,绕过大部分反爬机制,稳定性更高:

  1. 安装依赖:
    pip install playwright
    playwright install chrome
    
  2. 示例代码:
    from playwright.sync_api import sync_playwright
    from bs4 import BeautifulSoup
    
    with sync_playwright() as p:
        # 启动浏览器(headless=False可看到浏览器窗口,调试用;上线可改为True)
        browser = p.chromium.launch(headless=False)
        page = browser.new_page()
        # 访问目标页面
        page.goto('https://m.happymh.com/manga/woshidashenxian')
        # 等待页面加载完成
        page.wait_for_load_state('networkidle')
        # 获取页面源码
        html_content = page.content()
        # 解析内容
        soup = BeautifulSoup(html_content, 'html.parser')
        print(soup.title.text)
        # 关闭浏览器
        browser.close()
    

方案3:精简请求头(辅助优化)

如果坚持用原生requests,先移除失效的cookie和限制字段,保留必要请求头尝试:

from bs4 import BeautifulSoup 
import requests

headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.102 Safari/537.36',
    'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'accept-language': 'en-US,en;q=0.9'
}

link = 'https://m.happymh.com/manga/woshidashenxian'
req = requests.get(link, headers=headers)
print(req.status_code)

注意:这种方法大概率还是会被Cloudflare拦截,仅作基础尝试。

内容的提问来源于stack exchange,提问作者Mohamad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 08:21:57