You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取医生页面遇ElementClickInterceptedException翻页失败求助

Selenium爬取医生链接时,翻页按钮被广告iframe拦截无法点击

问题描述

我用Selenium爬取https://health.usnews.com/doctors/new-jersey页面的医生主页href链接,提取和去重功能正常,但翻页逻辑完全失效——要么程序陷入循环,要么直接退出。点击「Next Page」按钮时触发ElementClickInterceptedException,提示按钮被广告iframe遮挡。

错误信息

ElementClickInterceptedException: Message: element click intercepted: Element <button title="Next Page" data-page="2" class="page-numbers__Button-sc-138ov1k-4 page-numbers__NextButton-sc-138ov1k-6 
biGwcb hoTHNb"></button> is not clickable at point (456, 865). Other element would receive the click: <iframe frameborder="0" src="https://8305b9cc4359dda504fe47652a5a9af6.safeframe.googlesyndication.com/safeframe/1-0-40/html/container.html" id="google_ads_iframe_/4020/usn.health/healthcare/doctors/search/newjersey_1" title="3rd party ad content" name="" scrolling="no" marginwidth="0" marginheight="0" width="320" height="50" data-is-safeframe="true" sandbox="allow-forms allow-popups allow-popups-to-escape-sandbox allow-same-origin allow-scripts allow-top-navigation-by-user-activation" aria-label="Advertisement" tabindex="0" data-google-container-id="2" style="border: 0px; vertical-align: bottom;" data-load-complete="true"></iframe>
  (Session info: chrome=122.0.6261.70)

同时出现以下日志错误:

[10620:16572:0228/204556.116:ERROR:device_event_log_impl.cc(192)] [20:45:56.116] USB: usb_service_win.cc:105 SetupDiGetDeviceProperty({{A45C254E-DF1C-4EFD-8020-67D146A850E0}, 6}) failed: Element not found. (0x490)
[11692:2400:0228/204600.897:ERROR:ssl_client_socket_impl.cc(970)] handshake failed; returned -1, SSL error code 1, net_error -202

现有翻页逻辑代码

def getPageHLinks(url, count, tempSet, driver):

    junkLinks = ['https://www.usnews.com/features/info/privacy/state-notice',
                 'https://www.usnews.com/features/info/privacy'
                 ]
    try:
        # Extract the page's HTML
        doctors_HLinks_Set = tempSet  # Place hrefs into a set to remove duplicates
        doctors_HLinks = driver.find_elements(By.XPATH, "//a[@class='Anchor-byh49a-0 eMEqFO']")

        for doctors_HLink in doctors_HLinks:
            if doctors_HLink in junkLinks:  # Removal clause for a useless link
                continue
            doctors_HLinks_Set.add(doctors_HLink.get_attribute('href'))

        print(doctors_HLinks_Set)

       
        if count != 51:  # Page iteration automation, starts at page 2
            wait = WebDriverWait(driver, 10)

            try:
                button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title = 'Next Page' and @class= 'page-numbers__Button-sc-138ov1k-4 page-numbers__NextButton-sc-138ov1k-6 biGwcb hoTHNb']")))
                ActionChains(driver).move_to_element(button).perform()

                time.sleep(5)

                button.click() # Click the button 
                count += 1
                return getPageHLinks(url, count, doctors_HLinks_Set, driver)
            
            except ElementClickInterceptedException as e:
                print(f"ElementClickInterceptedException: {e}")
                driver.quit()
                return doctors_HLinks_Set
            
        else:  # Reached page 50, search is complete so output the data
            driver.quit()
            return doctors_HLinks_Set
        
    except NoSuchWindowException as e:
        print(f"No such window: {e}")
        driver.quit()
        return tempSet

已尝试的无效方案

  • 添加等待时间等待广告超时
  • 为ChromeDriver安装广告拦截扩展
  • 使用ActionChain滚动到页面底部后点击按钮

可行解决方案

1. 用JavaScript直接触发点击(最推荐)

元素被遮挡时,JS点击可以绕过视觉层的拦截,直接调用元素的点击事件:

button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']")))
# 替换原button.click()为下面这句
driver.execute_script("arguments[0].click();", button)

2. 隐藏广告iframe

先定位广告iframe,用JS将其隐藏,再点击按钮:

# 隐藏遮挡的广告iframe
driver.execute_script("document.getElementById('google_ads_iframe_/4020/usn.health/healthcare/doctors/search/newjersey_1').style.display = 'none';")
# 再执行点击
button.click()

3. 优化元素定位规则

原代码依赖完整的动态class名称,这类class容易随网站更新变化,建议简化定位:

# 用contains匹配部分class名称,提升兼容性
button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page' and contains(@class, 'NextButton')]")))

4. 调整滚动位置确保按钮完全可见

滚动到按钮所在位置并居中,避免被页面元素遮挡:

button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']")))
# 滚动到按钮居中位置
driver.execute_script("arguments[0].scrollIntoView({block: 'center'});")
time.sleep(1) # 留1秒等待滚动完成
button.click()

5. 替换递归为循环(解决循环/退出问题)

原代码用递归处理翻页,容易触发栈溢出或逻辑循环,改成循环结构更稳定:

def getPageHLinks(url, driver):
    junkLinks = {'https://www.usnews.com/features/info/privacy/state-notice',
                 'https://www.usnews.com/features/info/privacy'}
    doctors_HLinks_Set = set()
    count = 1
    max_pages = 50  # 爬取到第50页停止

    while count <= max_pages:
        # 提取当前页所有医生链接
        doctors_HLinks = driver.find_elements(By.XPATH, "//a[@class='Anchor-byh49a-0 eMEqFO']")
        for link in doctors_HLinks:
            href = link.get_attribute('href')
            if href and href not in junkLinks:
                doctors_HLinks_Set.add(href)
        
        print(f"已完成第{count}页爬取,当前累计链接数:{len(doctors_HLinks_Set)}")

        # 最后一页不需要翻页
        if count == max_pages:
            break

        # 执行翻页操作
        try:
            wait = WebDriverWait(driver, 10)
            button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']")))
            # 用JS点击绕过遮挡
            driver.execute_script("arguments[0].click();", button)
            # 等待页面加载完成(等待旧页第一个元素失效)
            wait.until(EC.staleness_of(doctors_HLinks[0]))
            count += 1
        except Exception as e:
            print(f"第{count+1}页翻页失败:{str(e)}")
            break

    driver.quit()
    return doctors_HLinks_Set

内容的提问来源于stack exchange,提问作者Michael Craig

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 22:54:53