如何在Selenium WebDriver中绕过reCAPTCHA实现谷歌搜索链接爬取?
问题描述
我尝试爬取谷歌搜索页面中的新闻链接,以下是搜索结果示例:
27 Dec 2017 – 24 Feb 2018 All results Clear Page 2 of 16 results (0.38 seconds) Nirav Modi, the jewellery designer at the centre of Rs ... businesstoday.in https://www.businesstoday.in › LATEST › Corporate 15-Feb-2018 — Union Bank of India, Allahabad Bank and Axis Bank are reported to have offered credit based on letters of undertaking (LOUs) issued by PNB. Published on: Feb 15 ... Auto Expo 2018: Tata Motors unveils H5X SUV, 45X ... businesstoday.in https://www.businesstoday.in › Auto 07-Feb-2018 — Nifty Bank can see significant upside; ICICI Bank, Axis Bank, SBI, BoB and Canara Bank may do well. RECOMMENDED. New vs old Parliament building: Government ... Implications of 10% long term capital gains tax proposed in ... businesstoday.in https://www.businesstoday.in › ... › Columns 03-Feb-2018 — The Finance Bill 2018 has further proposed significant changes in these taxation provisions relating to long term capital gain on transfer of equity shares ...
爬取链接的代码如下:
driver = webdriver.Chrome('E:\\Business_today\\chromedriver.exe') driver.get(urls[3]) time.sleep(10) link = set() i = 1 while(True): if i%10 == 0: print('pages done'+str(i)) try: myDiv = driver.find_element(By.CLASS_NAME,'v7W49e') t=myDiv.get_attribute("outerHTML") html_soup_object = BeautifulSoup(t, 'html.parser') find_element2 = html_soup_object.find_all('div') for ele in find_element2: for a_elm in ele.find_all("a"): # print(a_elm.attrs["href"]) if a_elm.attrs["href"].startswith('https://www.livemint.com/'): link.add(a_elm.attrs["href"]) time.sleep(5) try: link_element = driver.find_element(By.ID,'pnnext') link_element.click() i = i+ 1 except NoSuchElementException: break except NoSuchElementException: break driver.quit() links = list(link) # res2 = [links[i] for i in range(len(links)) if i % 2 != 0] print('total links ='+str(len(links))) total += links print("total home_page links = "+str(len(total))) time.sleep(10)
其中用于跳转谷歌搜索下一页的代码块:
try: link_element = driver.find_element(By.ID,'pnnext') link_element.click()
但进行多次搜索后,页面出现reCAPTCHA验证,导致WebDriver停止运行,请问是否有办法绕过该reCAPTCHA验证?
解决建议
- 模拟真实用户行为:
- 不要固定
time.sleep(5),改用随机间隔(比如2-8秒),避免机械性的请求节奏。 - 加入鼠标随机移动、滚动页面等操作,模拟用户浏览行为,而非直接点击下一页。
- 不要固定
- 优化请求特征:
- 启动Chrome时添加参数禁用自动化检测、模拟真实浏览器UA:
options = webdriver.ChromeOptions() options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome('E:\\Business_today\\chromedriver.exe', options=options) - 启用浏览器缓存和Cookie,让请求更接近长期使用的真实用户。
- 启动Chrome时添加参数禁用自动化检测、模拟真实浏览器UA:
- 限制请求频率:
- 减少单次爬取的页面数量,分批次执行,比如每爬10页就暂停10-15分钟再继续。
- 轮换代理IP:
- 单一IP频繁请求易被标记,使用高匿代理池轮换IP,避免触发验证机制。
- 优先使用官方API:
- 谷歌提供了Custom Search JSON API,通过正规渠道获取搜索结果,完全不会遇到reCAPTCHA问题,虽有请求额度限制,但稳定性远高于爬虫。
- 临时手动处理:
- 遇到验证时,添加等待逻辑让用户手动完成验证后再继续:
try: link_element = driver.find_element(By.ID,'pnnext') link_element.click() except (NoSuchElementException, ElementClickInterceptedException): input("完成reCAPTCHA验证后按回车继续...") link_element = driver.find_element(By.ID,'pnnext') link_element.click()
- 遇到验证时,添加等待逻辑让用户手动完成验证后再继续:
注意:绕过reCAPTCHA可能违反谷歌服务条款,爬取时请确保符合目标网站规则及相关法律法规。
内容的提问来源于stack exchange,提问作者pratik patil
相关产品推荐
相关产品推荐

