使用Puppeteer时如何避免Chrome的即时机器人拦截?
Google搜索拦截Puppeteer浏览器的原因及解决办法
核心原因
Google通过多维度浏览器指纹和行为特征识别自动化工具,即使是非无头模式,Chrome for Testing和Puppeteer的默认配置仍会留下明显自动化痕迹:
- Chrome for Testing专属标识:该版本内置测试用特殊标记,Google可直接识别。
- 指纹不一致:你生成的User-Agent版本(134.x)与HTTP头中
sec-ch-ua的版本(119.x)不匹配,触发校验失败。 - 启动参数特征:
--no-sandbox、--disable-gpu等参数在普通Chrome中不会默认启用,成为识别标记。 - 自动化属性残留:即使添加了
--disable-blink-features=AutomationControlled,仍可能存在未覆盖的自动化相关属性(如window.navigator.webdriver)。 - 无真实用户数据:新创建的
userDataDir没有历史记录、Cookie和登录状态,被判定为异常新用户。
针对性解决步骤
1. 替换为普通Chrome浏览器执行路径
Chrome for Testing的测试标识是主要触发点,直接指定本地普通Chrome的路径:
const browser = await puppeteerExtra.launch({ headless: false, executablePath: '/Applications/Google Chrome.app/Contents/MacOS/Google Chrome', // macOS示例,Windows/Linux替换为对应路径 userDataDir: userDataDir, // 建议使用普通Chrome的用户目录,比如~/Library/Application Support/Google/Chrome/Default(macOS) args: [ '--disable-blink-features=AutomationControlled', '--start-maximized', '--disable-dev-shm-usage', '--no-first-run', '--no-default-browser-check', '--disable-extensions-except=', // 禁用所有扩展避免干扰 '--disable-plugins', ], });
2. 修正HTTP头与User-Agent的一致性
修改generateRealisticHTTPHeaders函数,让sec-ch-ua版本与生成的Chrome版本同步:
async function generateRealisticHTTPHeaders(userAgent) { // 从User-Agent中提取Chrome版本号 const chromeVersion = userAgent.match(/Chrome\/(\d+\.\d+\.\d+\.\d+)/)[1]; const majorVersion = chromeVersion.split('.')[0]; // 提取主版本号 return { accept: "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8", "accept-language": "en-US,en;q=0.9", "accept-encoding": "gzip, deflate, br", "cache-control": "max-age=0", "sec-ch-ua": `"Google Chrome";v="${majorVersion}", "Chromium";v="${majorVersion}", "Not?A_Brand";v="24"`, "sec-ch-ua-mobile": "?0", "sec-ch-ua-platform": os.platform() === "darwin" ? '"macOS"' : os.platform() === "win32" ? '"Windows"' : '"Linux"', "sec-fetch-dest": "document", "sec-fetch-mode": "navigate", "sec-fetch-site": "none", "sec-fetch-user": "?1", "upgrade-insecure-requests": "1", "user-agent": userAgent, }; }
3. 正确配置Stealth插件
必须在launch之前注册Stealth插件,确保生效:
const puppeteerExtra = require('puppeteer-extra'); const StealthPlugin = require('puppeteer-extra-plugin-stealth'); // 注册插件要在启动浏览器之前 puppeteerExtra.use(StealthPlugin()); // 再启动浏览器 const browser = await puppeteerExtra.launch({...});
4. 移除冗余的历史伪造代码
你的injectFakeHistory函数可能触发异常的history行为,反而被检测。直接复用普通Chrome的用户目录,让浏览器拥有真实历史记录和Cookie,比伪造更有效。如果必须用新目录,建议先手动访问几个Google服务(如Gmail、Maps)再搜索。
5. 添加真实交互模拟
即使手动操作,也可以模拟真实用户行为降低检测概率:
// 打开Google首页后随机等待1-3秒 await page.goto('https://www.google.com'); await page.waitForTimeout(Math.random() * 2000 + 1000); // 模拟鼠标移动到搜索框并点击 await page.mouse.move(Math.random() * 100 + 50, Math.random() * 50 + 50); await page.click('input[name="q"]'); // 逐字符输入搜索关键词(模拟真实打字) const searchQuery = 'your search query'; for (const char of searchQuery) { await page.keyboard.type(char); await page.waitForTimeout(Math.random() * 100 + 50); // 每个字符间隔50-150ms } await page.keyboard.press('Enter');
6. 手动兜底覆盖自动化属性
即使使用Stealth插件,也可以添加兜底代码确保所有自动化属性被覆盖:
await page.evaluateOnNewDocument(() => { Object.defineProperty(navigator, 'webdriver', { get: () => undefined, }); if (window.chrome) { window.chrome.runtime = {}; } });
总结
Google的检测是多维度的,单一手段无法规避。核心思路是让Puppeteer启动的浏览器尽可能接近普通用户的Chrome环境:使用真实Chrome路径、保持指纹一致、复用真实用户数据、模拟真实交互行为。
内容的提问来源于stack exchange,提问作者Nate
相关产品推荐
相关产品推荐

