使用Selenium爬取数据时遭遇NoAlertPresentException错误求助
解决Selenium爬取时的NoAlertPresentException错误
问题描述
爬取https://mspotrace.org.my/Sccs_list时频繁出现NoAlertPresentException,错误信息如下:
NoAlertPresentException: no alert open (Session info: chrome=109.0.5414.121) (Driver info: chromedriver=2.35.528161 (5b82f2d2aae0ca24b877009200ced9065a772e73),platform=Windows NT 10.0.19045 x86_64)
调整time.sleep时长后问题仍未解决,原代码如下:
from selenium import webdriver from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time driver = webdriver.Chrome() driver.get('https://mspotrace.org.my/Sccs_list') time.sleep(20) # Get list of elements elements = WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']"))) # Loop through element popups and pull details of facilities into DF pos = 0 df = pd.DataFrame(columns=['facility_name','other_details','gmaps_url']) df_out = pd.DataFrame(columns=['facility_name','other_details','gmaps_url']) for iii in range(1,10): # testing with 10 pages for element in elements: try: data = [] element.click() time.sleep(10) facility_name = driver.find_element_by_xpath('//h4[@class="modal-title"]').text other_details = driver.find_element_by_xpath('//div[@class="modal-body"]').text map_url = driver.find_element_by_xpath("//a[contains(@href,'https://maps.google.com/maps?ll=')]") gmaps_url = str(map_url.get_attribute('href')) time.sleep(20) data.append(facility_name) data.append(other_details) data.append(gmaps_url) df.loc[pos] = data WebDriverWait(driver,1).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label='Close'] > span"))).click() # close popup window print("Scraping info for",facility_name,"") time.sleep(10) pos+=1 except Exception: alert = driver.switch_to.alert print("No geo location information") alert.accept() pass # click next btnNext = driver.find_element(By.XPATH,'//*[@id="dTable_next"]/a') driver.execute_script("arguments[0].scrollIntoView();", btnNext) driver.execute_script("arguments[0].click();", btnNext) time.sleep(5) # create outputs in df_out df_out = df_out.append(df) # Get list of elements again elements = WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']"))) # Resetting vars again pos = 0 df = pd.DataFrame(columns=['facility_name','other_details','gmaps_url'])
解决方案
1. 修复异常捕获逻辑,避免盲目切换Alert
原代码用except Exception:捕获所有异常,随后直接尝试切换Alert,导致无Alert时触发NoAlertPresentException。需分场景捕获特定异常:
- 先捕获元素查找失败的异常(比如找不到地图链接时可能触发Alert)
- 单独捕获
NoAlertPresentException,避免无效切换
2. 用显式等待判断Alert是否存在
不要直接调用driver.switch_to.alert,而是用WebDriverWait等待Alert出现,超时则判定无Alert:
from selenium.common.exceptions import NoAlertPresentException, TimeoutException try: alert = WebDriverWait(driver, 3).until(EC.alert_is_present()) alert.accept() print("No geo location information") except (TimeoutException, NoAlertPresentException): pass
3. 匹配Chrome与ChromeDriver版本
你的Chrome版本是109,但ChromeDriver是2.35,版本严重不兼容,这会导致包括Alert识别在内的各种异常。下载对应Chrome版本的ChromeDriver(Chrome 109对应ChromeDriver 109.x系列),确保版本一致。
4. 替换过时API,用显式等待替代固定sleep
原代码中find_element_by_xpath是过时API,改用find_element(By.XPATH, ...);同时减少time.sleep,用显式等待等待元素加载,提升稳定性:
比如点击元素后等待modal标题出现:
element.click() # 等待modal标题加载 facility_name = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, '//h4[@class="modal-title"]')) ).text
修改后的完整代码
from selenium import webdriver from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoAlertPresentException, TimeoutException, NoSuchElementException import pandas as pd driver = webdriver.Chrome() driver.get('https://mspotrace.org.my/Sccs_list') # 等待页面加载完成 WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']"))) df_out = pd.DataFrame(columns=['facility_name','other_details','gmaps_url']) for iii in range(1,10): # testing with 10 pages # 重新获取当前页的元素列表 elements = WebDriverWait(driver, 20).until(EC.presence_of_all_elements_located((By.XPATH, "//a[@title='View on Map']"))) df = pd.DataFrame(columns=['facility_name','other_details','gmaps_url']) pos = 0 for element in elements: try: data = [] element.click() # 等待modal标题可见 facility_name = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, '//h4[@class="modal-title"]')) ).text # 等待modal内容加载 other_details = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, '//div[@class="modal-body"]')) ).text try: # 尝试获取地图链接 map_url = WebDriverWait(driver, 5).until( EC.presence_of_element_located((By.XPATH, "//a[contains(@href,'https://maps.google.com/maps?ll=')]")) ) gmaps_url = map_url.get_attribute('href') except (TimeoutException, NoSuchElementException): # 找不到地图链接时检查Alert try: alert = WebDriverWait(driver, 3).until(EC.alert_is_present()) alert.accept() print(f"No geo location information for {facility_name}") gmaps_url = None except (TimeoutException, NoAlertPresentException): gmaps_url = None data.extend([facility_name, other_details, gmaps_url]) df.loc[pos] = data # 关闭modal close_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "button[aria-label='Close'] > span")) ) close_btn.click() print(f"Scraping info for {facility_name}") pos += 1 except Exception as e: print(f"Error processing element: {str(e)}") # 尝试关闭可能存在的modal或Alert try: alert = driver.switch_to.alert alert.accept() except NoAlertPresentException: pass try: close_btn = driver.find_element(By.CSS_SELECTOR, "button[aria-label='Close'] > span") if close_btn.is_displayed(): close_btn.click() except: pass continue # 点击下一页 try: btnNext = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="dTable_next"]/a')) ) driver.execute_script("arguments[0].scrollIntoView();", btnNext) btnNext.click() # 等待下一页加载完成 WebDriverWait(driver, 10).until( EC.staleness_of(elements[0]) # 等待上一页元素失效,说明页面已刷新 ) except Exception as e: print(f"Failed to click next page: {str(e)}") break # 合并数据 df_out = pd.concat([df_out, df], ignore_index=True) # 保存结果 df_out.to_csv('facilities.csv', index=False) driver.quit()
内容的提问来源于stack exchange,提问作者Funkeh-Monkeh
相关产品推荐
相关产品推荐

