You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Puppeteer时如何避免Chrome的即时机器人拦截?

Google搜索拦截Puppeteer浏览器的原因及解决办法

核心原因

Google通过多维度浏览器指纹和行为特征识别自动化工具,即使是非无头模式,Chrome for Testing和Puppeteer的默认配置仍会留下明显自动化痕迹:

  • Chrome for Testing专属标识:该版本内置测试用特殊标记,Google可直接识别。
  • 指纹不一致:你生成的User-Agent版本(134.x)与HTTP头中sec-ch-ua的版本(119.x)不匹配,触发校验失败。
  • 启动参数特征:--no-sandbox、--disable-gpu等参数在普通Chrome中不会默认启用,成为识别标记。
  • 自动化属性残留:即使添加了--disable-blink-features=AutomationControlled,仍可能存在未覆盖的自动化相关属性(如window.navigator.webdriver)。
  • 无真实用户数据:新创建的userDataDir没有历史记录、Cookie和登录状态,被判定为异常新用户。

针对性解决步骤

1. 替换为普通Chrome浏览器执行路径

Chrome for Testing的测试标识是主要触发点,直接指定本地普通Chrome的路径:

const browser = await puppeteerExtra.launch({
  headless: false,
  executablePath: '/Applications/Google Chrome.app/Contents/MacOS/Google Chrome', // macOS示例,Windows/Linux替换为对应路径
  userDataDir: userDataDir, // 建议使用普通Chrome的用户目录,比如~/Library/Application Support/Google/Chrome/Default(macOS)
  args: [
    '--disable-blink-features=AutomationControlled',
    '--start-maximized',
    '--disable-dev-shm-usage',
    '--no-first-run',
    '--no-default-browser-check',
    '--disable-extensions-except=', // 禁用所有扩展避免干扰
    '--disable-plugins',
  ],
});

2. 修正HTTP头与User-Agent的一致性

修改generateRealisticHTTPHeaders函数,让sec-ch-ua版本与生成的Chrome版本同步:

async function generateRealisticHTTPHeaders(userAgent) {
  // 从User-Agent中提取Chrome版本号
  const chromeVersion = userAgent.match(/Chrome\/(\d+\.\d+\.\d+\.\d+)/)[1];
  const majorVersion = chromeVersion.split('.')[0]; // 提取主版本号

  return {
    accept: "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8",
    "accept-language": "en-US,en;q=0.9",
    "accept-encoding": "gzip, deflate, br",
    "cache-control": "max-age=0",
    "sec-ch-ua": `"Google Chrome";v="${majorVersion}", "Chromium";v="${majorVersion}", "Not?A_Brand";v="24"`,
    "sec-ch-ua-mobile": "?0",
    "sec-ch-ua-platform": os.platform() === "darwin" ? '"macOS"' : os.platform() === "win32" ? '"Windows"' : '"Linux"',
    "sec-fetch-dest": "document",
    "sec-fetch-mode": "navigate",
    "sec-fetch-site": "none",
    "sec-fetch-user": "?1",
    "upgrade-insecure-requests": "1",
    "user-agent": userAgent,
  };
}

3. 正确配置Stealth插件

必须在launch之前注册Stealth插件,确保生效:

const puppeteerExtra = require('puppeteer-extra');
const StealthPlugin = require('puppeteer-extra-plugin-stealth');

// 注册插件要在启动浏览器之前
puppeteerExtra.use(StealthPlugin());

// 再启动浏览器
const browser = await puppeteerExtra.launch({...});

4. 移除冗余的历史伪造代码

你的injectFakeHistory函数可能触发异常的history行为,反而被检测。直接复用普通Chrome的用户目录,让浏览器拥有真实历史记录和Cookie,比伪造更有效。如果必须用新目录,建议先手动访问几个Google服务(如Gmail、Maps)再搜索。

5. 添加真实交互模拟

即使手动操作,也可以模拟真实用户行为降低检测概率:

// 打开Google首页后随机等待1-3秒
await page.goto('https://www.google.com');
await page.waitForTimeout(Math.random() * 2000 + 1000);

// 模拟鼠标移动到搜索框并点击
await page.mouse.move(Math.random() * 100 + 50, Math.random() * 50 + 50);
await page.click('input[name="q"]');

// 逐字符输入搜索关键词(模拟真实打字)
const searchQuery = 'your search query';
for (const char of searchQuery) {
  await page.keyboard.type(char);
  await page.waitForTimeout(Math.random() * 100 + 50); // 每个字符间隔50-150ms
}
await page.keyboard.press('Enter');

6. 手动兜底覆盖自动化属性

即使使用Stealth插件,也可以添加兜底代码确保所有自动化属性被覆盖:

await page.evaluateOnNewDocument(() => {
  Object.defineProperty(navigator, 'webdriver', {
    get: () => undefined,
  });

  if (window.chrome) {
    window.chrome.runtime = {};
  }
});

总结

Google的检测是多维度的,单一手段无法规避。核心思路是让Puppeteer启动的浏览器尽可能接近普通用户的Chrome环境:使用真实Chrome路径、保持指纹一致、复用真实用户数据、模拟真实交互行为。

内容的提问来源于stack exchange,提问作者Nate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 22:35:54