网页爬取时如何绕过Captcha?Selenium爬取车辆详情遭人机验证
Hey, I’ve run into exactly this issue with car listing platforms like Autoscout24—their anti-scraping systems are tuned to flag repetitive, non-human browsing patterns. Here are some proven fixes to avoid hitting that verification wall every 30 pages:
Add random, human-like delays
driver.implicitly_wait(20)only waits for elements to load—it doesn’t prevent anti-scraping triggers. You need to add random pauses between actions (like clicking "next page") to mimic how a real user would browse. Use something like:import random import time # After each page navigation time.sleep(random.uniform(2, 6)) # Random wait between 2-6 secondsAvoid fixed delays—bots use consistent timings, while humans vary their pace.
Mimic real browsing behavior
Autoscout24 tracks more than just request frequency. Try adding these actions:- Randomly scroll up/down the page before navigating to the next one (use
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")and random scroll positions) - Occasionally click on a random car listing, spend a few seconds on it, then go back to the list
- Use Selenium’s
ActionChainsto simulate mouse movements across the page, not just direct clicks
- Randomly scroll up/down the page before navigating to the next one (use
Rotate user agents
Selenium’s default user agent is easy to spot. Spoof a real browser’s user agent and rotate it periodically. You can find valid desktop user agents online, then set them via ChromeOptions:from selenium.webdriver.chrome.options import Options options = Options() user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.5 Safari/605.1.15", # Add more real user agents here ] options.add_argument(f'user-agent={random.choice(user_agents)}') driver = webdriver.Chrome(options=options)Use residential proxies
If you’re using the same IP for all requests, Autoscout24 will flag it quickly. Residential proxies use real home IP addresses, making your traffic look like a genuine user. Rotate proxies every 10-20 pages to avoid detection.Optimize headless mode (if you use it)
Headless Chrome is easier to detect. If you must use it, add these flags to make it look more like a regular browser:options.add_argument("--headless=new") options.add_argument("--window-size=1920,1080") options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False)Alternatively, skip headless mode entirely if possible—regular browser windows are harder to flag.
Try undetected-chromedriver
This is a modified version of ChromeDriver that bypasses most Selenium detection checks. It patches the browser to hide automation signals that Autoscout24 looks for. Install it via pip (pip install undetected-chromedriver) and replace your regular driver initialization with it.Limit your crawl rate
Even with all the above, pushing too hard in a short time will trigger verification. Split your crawl into batches—e.g., crawl 25 pages, then pause for 10-15 minutes, then resume. Or spread your scraping over hours/days instead of trying to grab everything at once.
内容的提问来源于stack exchange,提问作者Michaella

