You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Selenium WebDriver中绕过reCAPTCHA实现谷歌搜索链接爬取?

问题描述

我尝试爬取谷歌搜索页面中的新闻链接,以下是搜索结果示例:

27 Dec 2017 – 24 Feb 2018
All results
Clear
Page 2 of 16 results (0.38 seconds) 

Nirav Modi, the jewellery designer at the centre of Rs ...

businesstoday.in
https://www.businesstoday.in › LATEST › Corporate
15-Feb-2018 — Union Bank of India, Allahabad Bank and Axis Bank are reported to have offered credit based on letters of undertaking (LOUs) issued by PNB. Published on: Feb 15 ...

Auto Expo 2018: Tata Motors unveils H5X SUV, 45X ...

businesstoday.in
https://www.businesstoday.in › Auto
07-Feb-2018 — Nifty Bank can see significant upside; ICICI Bank, Axis Bank, SBI, BoB and Canara Bank may do well. RECOMMENDED. New vs old Parliament building: Government ...

Implications of 10% long term capital gains tax proposed in ...

businesstoday.in
https://www.businesstoday.in › ... › Columns
03-Feb-2018 — The Finance Bill 2018 has further proposed significant changes in these taxation provisions relating to long term capital gain on transfer of equity shares ...

爬取链接的代码如下:

driver = webdriver.Chrome('E:\\Business_today\\chromedriver.exe')
driver.get(urls[3])
time.sleep(10)

link = set()
i = 1
while(True):
    if i%10 == 0:
        print('pages done'+str(i))
    try:
        myDiv = driver.find_element(By.CLASS_NAME,'v7W49e')
        t=myDiv.get_attribute("outerHTML")
        html_soup_object = BeautifulSoup(t, 'html.parser')
        find_element2 = html_soup_object.find_all('div')

        for ele in find_element2:
            for a_elm in ele.find_all("a"):
                # print(a_elm.attrs["href"])
                if a_elm.attrs["href"].startswith('https://www.livemint.com/'):
                    link.add(a_elm.attrs["href"])
        time.sleep(5)
        try:
            link_element = driver.find_element(By.ID,'pnnext')
            link_element.click()  
            i = i+ 1
        except NoSuchElementException:
            break
    except NoSuchElementException:
        break
driver.quit()


links = list(link)
# res2 = [links[i] for i in range(len(links)) if i % 2 != 0]
print('total links ='+str(len(links)))
total += links
print("total home_page links = "+str(len(total)))
time.sleep(10)

其中用于跳转谷歌搜索下一页的代码块:

try:
    link_element = driver.find_element(By.ID,'pnnext')
    link_element.click()

但进行多次搜索后,页面出现reCAPTCHA验证,导致WebDriver停止运行,请问是否有办法绕过该reCAPTCHA验证?

解决建议
  • 模拟真实用户行为:
    • 不要固定time.sleep(5),改用随机间隔(比如2-8秒),避免机械性的请求节奏。
    • 加入鼠标随机移动、滚动页面等操作,模拟用户浏览行为,而非直接点击下一页。
  • 优化请求特征:
    • 启动Chrome时添加参数禁用自动化检测、模拟真实浏览器UA:
      options = webdriver.ChromeOptions()
      options.add_argument("--disable-blink-features=AutomationControlled")
      options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
      driver = webdriver.Chrome('E:\\Business_today\\chromedriver.exe', options=options)
      
    • 启用浏览器缓存和Cookie,让请求更接近长期使用的真实用户。
  • 限制请求频率:
    • 减少单次爬取的页面数量,分批次执行,比如每爬10页就暂停10-15分钟再继续。
  • 轮换代理IP:
    • 单一IP频繁请求易被标记,使用高匿代理池轮换IP,避免触发验证机制。
  • 优先使用官方API:
    • 谷歌提供了Custom Search JSON API,通过正规渠道获取搜索结果,完全不会遇到reCAPTCHA问题,虽有请求额度限制,但稳定性远高于爬虫。
  • 临时手动处理:
    • 遇到验证时,添加等待逻辑让用户手动完成验证后再继续:
      try:
          link_element = driver.find_element(By.ID,'pnnext')
          link_element.click()
      except (NoSuchElementException, ElementClickInterceptedException):
          input("完成reCAPTCHA验证后按回车继续...")
          link_element = driver.find_element(By.ID,'pnnext')
          link_element.click()
      

注意:绕过reCAPTCHA可能违反谷歌服务条款,爬取时请确保符合目标网站规则及相关法律法规。

内容的提问来源于stack exchange,提问作者pratik patil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 00:34:54