You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Puppeteer开发谷歌搜索爬虫遇reCAPTCHA,如何触发人工验证提示?

Handling reCAPTCHA with Manual Verification in Puppeteer

Hey Chase! Great job getting your crawler up and running, and kudos for doing the right thing by not trying to bypass reCAPTCHA. Let's fix this so you can manually complete the check when it pops up and resume your searches smoothly.

First, let's clear up why your current code isn't working: reCAPTCHA runs inside an <iframe>, so trying to select elements like #g-recaptcha-response directly from the main page won't work—those elements live inside the iframe's document, not the parent. But since you want to do manual verification anyway, we don't need to mess with those elements at all.

Here's a step-by-step solution tailored to your needs:

1. Improve reCAPTCHA Detection

Instead of just checking the content string, we can check for both Google's "unusual traffic" message and the presence of the reCAPTCHA iframe. This makes detection more reliable.

2. Pause for Manual Verification

When we spot a reCAPTCHA, we'll prompt you to complete it manually, then wait for your confirmation before proceeding. We'll use Node.js's built-in readline module to handle this input.

3. Reuse Browser Instance

Right now, you're launching a new browser for every search query—this is inefficient and more likely to trigger Google's anti-bot measures. Let's move the browser launch outside the loop to reuse it across all queries.

Modified Code Snippet

const fs = require('fs');
const puppeteer = require('puppeteer');
const cheerio = require('cheerio');
const LineByLine = require('line-by-line');
const readline = require('readline');

// Set up readline to get user input for verification prompts
const rl = readline.createInterface({
  input: process.stdin,
  output: process.stdout
});

// Helper function to wait for user to confirm they've completed reCAPTCHA
const waitForManualVerification = () => {
  return new Promise(resolve => {
    rl.question('👋 Please complete the reCAPTCHA in the browser, then press Enter to continue...', () => {
      resolve();
    });
  });
};

const lr = new LineByLine('dorks.txt', { skipEmptyLines: true });
const dorks = [];

lr.on('line', function(line) {
  dorks.push(line);
});

lr.on('end', async function () {
  console.log(`[*] Loaded ${dorks.length} dorks. Let's begin.`);
  
  // Launch browser ONCE, reuse for all queries to reduce bot detection
  const browser = await puppeteer.launch({ timeout: 10000, headless: false });
  
  for (let i = 0; i < dorks.length; i++) {
    const page = await browser.newPage();
    try {
      // Encode the query to handle special characters properly
      const encodedQuery = encodeURIComponent(dorks[i]);
      await page.goto(`https://www.google.com/search?num=100&q=${encodedQuery}`);
      
      // Check if reCAPTCHA or unusual traffic warning is present
      const hasCaptcha = await page.evaluate(() => {
        const hasUnusualTraffic = document.body.textContent.includes("Our systems have detected unusual traffic from your computer network.");
        const hasRecaptchaIframe = document.querySelector('iframe[src*="recaptcha"]') !== null;
        return hasUnusualTraffic || hasRecaptchaIframe;
      });
      
      if (hasCaptcha) {
        console.log(`[!] reCAPTCHA detected for query: ${dorks[i]}`);
        await waitForManualVerification();
        
        // Wait for navigation in case user clicks submit after verification
        try {
          await page.waitForNavigation({ timeout: 30000 });
        } catch (e) {
          // If no navigation (user already submitted), just proceed to extract content
          console.log("[*] No navigation detected, moving to content extraction.");
        }
      }
      
      // Extract and process page content
      const content = await page.content();
      const $ = cheerio.load(content);
      
      // Add your content processing logic here (e.g., scraping search results)
      console.log(`[*] Successfully processed query ${i+1}/${dorks.length}: ${dorks[i]}`);
      
    } catch (error) {
      console.error(`[!] Error processing query "${dorks[i]}":`, error.message);
    } finally {
      await page.close(); // Close the page after processing to free resources
    }
  }
  
  // Cleanup
  await browser.close();
  rl.close();
  console.log("[*] All queries processed! 🎉");
});

Key Changes Explained

  • Reused Browser Instance: Launches Puppeteer once before the loop, which is more efficient and less likely to trigger anti-bot checks.
  • Robust Captcha Detection: Uses page.evaluate() to check both the unusual traffic message and the reCAPTCHA iframe.
  • Manual Verification Flow: Pauses execution with a user-friendly prompt until you confirm you've completed the reCAPTCHA.
  • Encoded Query: Uses encodeURIComponent() to handle special characters in search queries (avoids broken URLs).
  • Graceful Error Handling: Catches navigation timeouts and other errors to keep the crawler running through issues.

Quick Note on PhantomJS

PhantomJS is no longer maintained, so Puppeteer is definitely the better choice here. It's actively supported, has excellent documentation, and integrates seamlessly with modern Node.js practices—no need to switch tools!

If you run into any snags with this code, or need help expanding the content-scraping part, feel free to ask.

内容的提问来源于stack exchange,提问作者Chase Taylor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:46:58