使用Puppeteer开发谷歌搜索爬虫遇reCAPTCHA,如何触发人工验证提示?
Hey Chase! Great job getting your crawler up and running, and kudos for doing the right thing by not trying to bypass reCAPTCHA. Let's fix this so you can manually complete the check when it pops up and resume your searches smoothly.
First, let's clear up why your current code isn't working: reCAPTCHA runs inside an <iframe>, so trying to select elements like #g-recaptcha-response directly from the main page won't work—those elements live inside the iframe's document, not the parent. But since you want to do manual verification anyway, we don't need to mess with those elements at all.
Here's a step-by-step solution tailored to your needs:
1. Improve reCAPTCHA Detection
Instead of just checking the content string, we can check for both Google's "unusual traffic" message and the presence of the reCAPTCHA iframe. This makes detection more reliable.
2. Pause for Manual Verification
When we spot a reCAPTCHA, we'll prompt you to complete it manually, then wait for your confirmation before proceeding. We'll use Node.js's built-in readline module to handle this input.
3. Reuse Browser Instance
Right now, you're launching a new browser for every search query—this is inefficient and more likely to trigger Google's anti-bot measures. Let's move the browser launch outside the loop to reuse it across all queries.
Modified Code Snippet
const fs = require('fs'); const puppeteer = require('puppeteer'); const cheerio = require('cheerio'); const LineByLine = require('line-by-line'); const readline = require('readline'); // Set up readline to get user input for verification prompts const rl = readline.createInterface({ input: process.stdin, output: process.stdout }); // Helper function to wait for user to confirm they've completed reCAPTCHA const waitForManualVerification = () => { return new Promise(resolve => { rl.question('👋 Please complete the reCAPTCHA in the browser, then press Enter to continue...', () => { resolve(); }); }); }; const lr = new LineByLine('dorks.txt', { skipEmptyLines: true }); const dorks = []; lr.on('line', function(line) { dorks.push(line); }); lr.on('end', async function () { console.log(`[*] Loaded ${dorks.length} dorks. Let's begin.`); // Launch browser ONCE, reuse for all queries to reduce bot detection const browser = await puppeteer.launch({ timeout: 10000, headless: false }); for (let i = 0; i < dorks.length; i++) { const page = await browser.newPage(); try { // Encode the query to handle special characters properly const encodedQuery = encodeURIComponent(dorks[i]); await page.goto(`https://www.google.com/search?num=100&q=${encodedQuery}`); // Check if reCAPTCHA or unusual traffic warning is present const hasCaptcha = await page.evaluate(() => { const hasUnusualTraffic = document.body.textContent.includes("Our systems have detected unusual traffic from your computer network."); const hasRecaptchaIframe = document.querySelector('iframe[src*="recaptcha"]') !== null; return hasUnusualTraffic || hasRecaptchaIframe; }); if (hasCaptcha) { console.log(`[!] reCAPTCHA detected for query: ${dorks[i]}`); await waitForManualVerification(); // Wait for navigation in case user clicks submit after verification try { await page.waitForNavigation({ timeout: 30000 }); } catch (e) { // If no navigation (user already submitted), just proceed to extract content console.log("[*] No navigation detected, moving to content extraction."); } } // Extract and process page content const content = await page.content(); const $ = cheerio.load(content); // Add your content processing logic here (e.g., scraping search results) console.log(`[*] Successfully processed query ${i+1}/${dorks.length}: ${dorks[i]}`); } catch (error) { console.error(`[!] Error processing query "${dorks[i]}":`, error.message); } finally { await page.close(); // Close the page after processing to free resources } } // Cleanup await browser.close(); rl.close(); console.log("[*] All queries processed! 🎉"); });
Key Changes Explained
- Reused Browser Instance: Launches Puppeteer once before the loop, which is more efficient and less likely to trigger anti-bot checks.
- Robust Captcha Detection: Uses
page.evaluate()to check both the unusual traffic message and the reCAPTCHA iframe. - Manual Verification Flow: Pauses execution with a user-friendly prompt until you confirm you've completed the reCAPTCHA.
- Encoded Query: Uses
encodeURIComponent()to handle special characters in search queries (avoids broken URLs). - Graceful Error Handling: Catches navigation timeouts and other errors to keep the crawler running through issues.
Quick Note on PhantomJS
PhantomJS is no longer maintained, so Puppeteer is definitely the better choice here. It's actively supported, has excellent documentation, and integrates seamlessly with modern Node.js practices—no need to switch tools!
If you run into any snags with this code, or need help expanding the content-scraping part, feel free to ask.
内容的提问来源于stack exchange,提问作者Chase Taylor

