网页爬取中绕过蜜罐(Honeypots)的技术方法咨询
Nice catch—this is a pretty crafty anti-scraping tactic built to fool basic scrapers that only pull raw HTML instead of rendering pages like a real browser. Let’s break down how it works and how you can get around it.
This is a classic "decoy element" trick, designed to separate human users (and real browsers) from automated scrapers:
- When you first send an HTTP GET request, the server returns a bunch of identical-looking button elements, all hidden by default (thanks to
style='display:none;'). Each button links to a unique URL with a differentcparameter—only one of these parameters is valid. - A real browser will load the page's CSS (or sometimes accompanying JavaScript) after fetching the HTML. There’s a hidden CSS rule (or JS snippet) that targets the specific div ID containing the valid button and changes its
displayproperty toblock(or another visible value), making it the only clickable option for humans. - The server checks if the
cparameter in your subsequent request matches the one from the visible button. Scrapers that just parse raw HTML will pick a random button or fail to identify the valid one, triggering a block.
To get past this, you need to replicate how a real browser processes the page—rendering CSS/JS and identifying the visible element. Here are your best options:
1. Use a Headless Browser (Most Reliable)
Headless tools like Puppeteer (Chrome-based) or Playwright can fully render pages, apply styles, and execute JavaScript exactly like a regular browser. Here’s a step-by-step approach:
- Launch the headless browser and navigate to your target page, making sure to carry over your trusted session cookies (from bypassing CAPTCHA).
- Wait for the page to fully load—this gives CSS/JS time to make the valid button visible.
- Query the DOM for the visible button (filter out elements with
display:none). - Extract the valid URL from the button’s
onclickattribute. - Use that URL to fetch your target data, reusing the same session to stay trusted.
Example Puppeteer snippet:
const puppeteer = require('puppeteer'); (async () => { // Launch browser with your existing session cookies (replace with your cookie data) const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.setCookie(...yourCaptchaBypassedCookies); await page.goto('YOUR_TARGET_PAGE_URL'); // Wait for the valid button to become visible const visibleButton = await page.waitForSelector('button.build:not([style*="display:none"])'); // Extract the valid URL from the onclick handler const onClickCode = await visibleButton.evaluate(el => el.getAttribute('onclick')); const validUrl = onClickCode.match(/window.location.href = '(.*)';/)[1]; // Navigate to the valid URL and extract your data await page.goto(validUrl); const targetData = await page.$eval('YOUR_DATA_SELECTOR', el => el.textContent); console.log(targetData); await browser.close(); })();
2. Parse CSS to Identify the Valid Button (Fragile but Possible)
If you don’t want to use a headless browser, you can:
- Fetch the page’s HTML and all linked CSS files.
- Search through the CSS for rules that target a specific div ID (like
#BA405352A9) and overridedisplay:nonewithdisplay:block(or similar). - Once you find the target ID, go back to the HTML and pull the button’s
onclickURL from that div.
Note: This only works if visibility is controlled by static CSS. If the site uses dynamic JavaScript to toggle visibility (e.g., adding/removing classes), this method will fail.
3. Inspect Network Traffic (Fallback)
Open your browser’s DevTools (Network tab), clear the cache, and reload the page. Watch for:
- AJAX requests that fetch the valid
cparameter separately. - JavaScript files that contain logic to reveal the correct button.
Sometimes the valid parameter is loaded dynamically, and the hidden buttons are just decoys. This is a good fallback if the first two methods hit roadblocks.
- Reuse your session: Don’t discard the cookies you used to bypass CAPTCHA—your session is trusted, and switching to a new cookie jar will trigger blocks.
- Mimic human timing: Add small delays between page navigation and clicks (e.g., 1-2 seconds) to avoid looking like a bot.
- Use a real User-Agent: Set your scraper’s User-Agent to match a modern browser (e.g., Chrome’s latest version) to avoid being flagged.
内容的提问来源于stack exchange,提问作者Xosrov

