You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取Ambition Box网站HTML代码的技术求助

问题

我用下面的Python代码请求AmbitionBox的公司列表页面:

# 模拟Chrome浏览器请求头
header = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36"
}
# 发起请求
response = requests.get("https://www.ambitionbox.com/list-of-companies?page=1", headers=header)
# 查看返回的页面源码片段
response.text[0:500]

预期能拿到正常的HTML片段:

<!doctype html>
<html data-n-head-ssr lang="en" data-n-head="%7B%22lang%22:%7B%22ssr%22:%22en%22%7D%7D">
  <head >
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width,initial-scale=1">
    <meta http-equiv="X-UA-Compatible" content="IE=edge"> 
    
    <script type="text/javascript">window.NREUM||(NREUM={}),NREUM.init={distributed_tracing:{enabled:!0}},window.NREUM||(NREUM={}),__nr_require=function(n,r,t){function o(e){if(!r[e]){var t=r[e]={exports:{}};n[e][0].call(t

但实际返回的是带refresh元标签的页面片段:

<!DOCTYPE html>
<html>
  <head>
    <meta charset="utf-8">
    <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no">
    <meta http-equiv="refresh" content="5; URL='/list-of-companies?page=1&amp;bm-verify=AAQAAAAH_____276GjTiSUrRZ57bXtQd2QOqDb6yzCSiw3txyMphpYFQNCJPEdUs0V5AOZD9bqWb9kr6Fp7sjO7l20QoJ9Br0-xif5TA0H7wrtpMZ81GeXv3aVnX5j-O5O3cZwHTOBn_dSsnlYEIfuZiiSi2pzTSogOOO9mPfZvC5y_WwYEtvkCBs7uPmtd4u17HOLEp6QC_7lIMkx77-pvpYjJ0FToKiIka8JU0IKmf95a5XjSIf_xu7Qv-GDvEwZqgJWfz

需要解决这个问题,获取目标网站的正常HTML代码。

解决方法

1. 处理Cloudflare基础验证(bm-verify参数)

返回的refresh标签里的bm-verify参数是Cloudflare的基础人机验证标识,直接提取该参数重新发起请求即可:

import requests
from bs4 import BeautifulSoup

header = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36"
}

# 第一次请求获取验证跳转URL
response = requests.get("https://www.ambitionbox.com/list-of-companies?page=1", headers=header)
soup = BeautifulSoup(response.text, "html.parser")
refresh_meta = soup.find("meta", {"http-equiv": "refresh"})

if refresh_meta:
    # 提取跳转路径并拼接完整URL
    refresh_path = refresh_meta["content"].split("URL=")[1].strip("'\"")
    full_url = f"https://www.ambitionbox.com{refresh_path}"
    # 第二次请求带验证参数的地址
    final_response = requests.get(full_url, headers=header)
    print(final_response.text[:500])

2. 补充完整请求头模拟真实浏览器

仅带User-Agent不足以绕过检测,补充常见请求头字段让请求更贴近真实浏览器行为:

header = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
    "Referer": "https://www.ambitionbox.com/",
    "DNT": "1",
    "Connection": "keep-alive",
    "Upgrade-Insecure-Requests": "1"
}

3. 使用会话保持Cookie状态

网站会通过Cookie识别会话,用requests.Session()自动维持Cookie,避免被判定为异常请求:

import requests

session = requests.Session()
header = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36",
    # 加上上述补充的请求头
}

# 先访问首页获取初始会话Cookie
session.get("https://www.ambitionbox.com/", headers=header)
# 再请求目标页面
response = session.get("https://www.ambitionbox.com/list-of-companies?page=1", headers=header)
print(response.text[:500])

4. 验证机制升级时用浏览器自动化工具

如果上述方法无效,说明网站启用了更严格的JS渲染或交互验证,用Selenium模拟真实浏览器操作:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36")
driver = webdriver.Chrome(options=options)

driver.get("https://www.ambitionbox.com/list-of-companies?page=1")
page_source = driver.page_source
print(page_source[:500])
driver.quit()

内容的提问来源于stack exchange,提问作者Gautam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 14:43:22