You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python requests绕过网站加载页面获取主页面内容?

解决requests获取网站主页面而非加载页面的问题

这种加载页面通常是网站通过前端延迟渲染、二次请求或请求头验证实现的,以下是仅用requests库的解决方案:

  • 模拟浏览器请求头
    很多网站会校验请求的User-Agent字段,判断是否为浏览器发起的请求。添加标准浏览器的UA头后,可能直接返回主页面:

    import requests
    
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    response = requests.get("https://example.com", headers=headers)
    print(response.text)
    
  • 直接请求主内容接口
    打开浏览器开发者工具的Network标签,观察加载页面后的异步请求,找到获取主内容的API接口,直接请求该接口(需带上必要的请求头和参数):

    import requests
    
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
        'X-Requested-With': 'XMLHttpRequest'  # 部分AJAX接口需此标识
    }
    # 替换为实际的主内容接口地址
    response = requests.get("https://example.com/api/main-content", headers=headers)
    print(response.json())  # 若返回JSON格式,根据实际情况调整解析方式
    
  • 利用Session保持Cookie
    部分网站会在第一次请求时设置Cookie,第二次请求才返回主页面。使用requests.Session可以自动保持Cookie:

    import requests
    
    session = requests.Session()
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    # 首次请求获取加载页面及Cookie
    session.get("https://example.com", headers=headers)
    # 二次请求获取主页面
    response = session.get("https://example.com", headers=headers)
    print(response.text)
    

若以上方法均无效,说明主内容完全依赖JavaScript动态渲染且无独立API,此时可能需要使用浏览器模拟工具,但优先尝试上述方案。

内容的提问来源于stack exchange,提问作者Anonymous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 02:37:25