Selenium爬取医生页面遇ElementClickInterceptedException翻页失败求助
问题描述
我用Selenium爬取https://health.usnews.com/doctors/new-jersey页面的医生主页href链接,提取和去重功能正常,但翻页逻辑完全失效——要么程序陷入循环,要么直接退出。点击「Next Page」按钮时触发ElementClickInterceptedException,提示按钮被广告iframe遮挡。
错误信息
ElementClickInterceptedException: Message: element click intercepted: Element <button title="Next Page" data-page="2" class="page-numbers__Button-sc-138ov1k-4 page-numbers__NextButton-sc-138ov1k-6 biGwcb hoTHNb"></button> is not clickable at point (456, 865). Other element would receive the click: <iframe frameborder="0" src="https://8305b9cc4359dda504fe47652a5a9af6.safeframe.googlesyndication.com/safeframe/1-0-40/html/container.html" id="google_ads_iframe_/4020/usn.health/healthcare/doctors/search/newjersey_1" title="3rd party ad content" name="" scrolling="no" marginwidth="0" marginheight="0" width="320" height="50" data-is-safeframe="true" sandbox="allow-forms allow-popups allow-popups-to-escape-sandbox allow-same-origin allow-scripts allow-top-navigation-by-user-activation" aria-label="Advertisement" tabindex="0" data-google-container-id="2" style="border: 0px; vertical-align: bottom;" data-load-complete="true"></iframe> (Session info: chrome=122.0.6261.70)
同时出现以下日志错误:
[10620:16572:0228/204556.116:ERROR:device_event_log_impl.cc(192)] [20:45:56.116] USB: usb_service_win.cc:105 SetupDiGetDeviceProperty({{A45C254E-DF1C-4EFD-8020-67D146A850E0}, 6}) failed: Element not found. (0x490)
[11692:2400:0228/204600.897:ERROR:ssl_client_socket_impl.cc(970)] handshake failed; returned -1, SSL error code 1, net_error -202
现有翻页逻辑代码
def getPageHLinks(url, count, tempSet, driver): junkLinks = ['https://www.usnews.com/features/info/privacy/state-notice', 'https://www.usnews.com/features/info/privacy' ] try: # Extract the page's HTML doctors_HLinks_Set = tempSet # Place hrefs into a set to remove duplicates doctors_HLinks = driver.find_elements(By.XPATH, "//a[@class='Anchor-byh49a-0 eMEqFO']") for doctors_HLink in doctors_HLinks: if doctors_HLink in junkLinks: # Removal clause for a useless link continue doctors_HLinks_Set.add(doctors_HLink.get_attribute('href')) print(doctors_HLinks_Set) if count != 51: # Page iteration automation, starts at page 2 wait = WebDriverWait(driver, 10) try: button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title = 'Next Page' and @class= 'page-numbers__Button-sc-138ov1k-4 page-numbers__NextButton-sc-138ov1k-6 biGwcb hoTHNb']"))) ActionChains(driver).move_to_element(button).perform() time.sleep(5) button.click() # Click the button count += 1 return getPageHLinks(url, count, doctors_HLinks_Set, driver) except ElementClickInterceptedException as e: print(f"ElementClickInterceptedException: {e}") driver.quit() return doctors_HLinks_Set else: # Reached page 50, search is complete so output the data driver.quit() return doctors_HLinks_Set except NoSuchWindowException as e: print(f"No such window: {e}") driver.quit() return tempSet
已尝试的无效方案
- 添加等待时间等待广告超时
- 为ChromeDriver安装广告拦截扩展
- 使用ActionChain滚动到页面底部后点击按钮
可行解决方案
1. 用JavaScript直接触发点击(最推荐)
元素被遮挡时,JS点击可以绕过视觉层的拦截,直接调用元素的点击事件:
button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']"))) # 替换原button.click()为下面这句 driver.execute_script("arguments[0].click();", button)
2. 隐藏广告iframe
先定位广告iframe,用JS将其隐藏,再点击按钮:
# 隐藏遮挡的广告iframe driver.execute_script("document.getElementById('google_ads_iframe_/4020/usn.health/healthcare/doctors/search/newjersey_1').style.display = 'none';") # 再执行点击 button.click()
3. 优化元素定位规则
原代码依赖完整的动态class名称,这类class容易随网站更新变化,建议简化定位:
# 用contains匹配部分class名称,提升兼容性 button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page' and contains(@class, 'NextButton')]")))
4. 调整滚动位置确保按钮完全可见
滚动到按钮所在位置并居中,避免被页面元素遮挡:
button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']"))) # 滚动到按钮居中位置 driver.execute_script("arguments[0].scrollIntoView({block: 'center'});") time.sleep(1) # 留1秒等待滚动完成 button.click()
5. 替换递归为循环(解决循环/退出问题)
原代码用递归处理翻页,容易触发栈溢出或逻辑循环,改成循环结构更稳定:
def getPageHLinks(url, driver): junkLinks = {'https://www.usnews.com/features/info/privacy/state-notice', 'https://www.usnews.com/features/info/privacy'} doctors_HLinks_Set = set() count = 1 max_pages = 50 # 爬取到第50页停止 while count <= max_pages: # 提取当前页所有医生链接 doctors_HLinks = driver.find_elements(By.XPATH, "//a[@class='Anchor-byh49a-0 eMEqFO']") for link in doctors_HLinks: href = link.get_attribute('href') if href and href not in junkLinks: doctors_HLinks_Set.add(href) print(f"已完成第{count}页爬取,当前累计链接数:{len(doctors_HLinks_Set)}") # 最后一页不需要翻页 if count == max_pages: break # 执行翻页操作 try: wait = WebDriverWait(driver, 10) button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@title='Next Page']"))) # 用JS点击绕过遮挡 driver.execute_script("arguments[0].click();", button) # 等待页面加载完成(等待旧页第一个元素失效) wait.until(EC.staleness_of(doctors_HLinks[0])) count += 1 except Exception as e: print(f"第{count+1}页翻页失败:{str(e)}") break driver.quit() return doctors_HLinks_Set
内容的提问来源于stack exchange,提问作者Michael Craig

