使用Selenium定位伦敦证交所页面元素随机失败,如何实现100%定位成功率?
解决伦敦证券交易所公告脚本随机定位失败问题
问题背景
运行脚本爬取伦敦证券交易所上市公司最新公告时,出现随机定位失败的情况:4次执行中有3次成功,偶尔找不到以下两个元素的XPath:
//*[@id="news-table-results"]/div[1]/form[1] //*[contains(text(), 'Show 500 news')]
已使用WebDriverWait等待元素加载,但仍无法保证100%成功率。
优化方案
以下是针对性的解决措施,可大幅提升元素定位的稳定性:
延长等待超时时间并更换等待条件
无头浏览器的加载速度通常比有头模式慢,原5秒超时可能不足以应对网络波动或页面渲染延迟。将超时时间延长至15秒,同时使用element_to_be_clickable替代visibility_of_element_located——因为需要点击元素,"可点击"状态比"可见"状态更能确保元素完全就绪。优化XPath定位表达式
原XPath依赖页面层级结构,一旦页面DOM微调就会失效。改用更稳定的属性组合定位:- 下拉框表单:从依赖固定层级的
//*[@id="news-table-results"]/div[1]/form[1],改为//form[contains(@class, 'js-show-results') and @action="#"],通过类名和属性组合定位,避免层级变化影响。 - "Show 500 news"选项:从
//*[contains(text(), 'Show 500 news')]改为//div[contains(@class, 'dropdown-menu')]//span[text()='Show 500 news'],限定在下拉菜单容器内精确匹配文本,防止误匹配其他区域的相似文本。
- 下拉框表单:从依赖固定层级的
增强无头浏览器的兼容性
给Chrome无头模式添加更多参数,模拟真实浏览器环境,避免被页面反爬机制拦截或渲染异常:chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")添加关键步骤的重试机制
针对容易失败的定位步骤,添加循环重试逻辑,处理偶发的网络或渲染异常:from selenium.common.exceptions import TimeoutException def retry_find_element(driver, by, value, max_retries=3, timeout=15): for _ in range(max_retries): try: return WebDriverWait(driver, timeout).until(EC.element_to_be_clickable((by, value))) except TimeoutException: driver.refresh() # 刷新页面后重试 raise TimeoutException(f"Failed to locate element after {max_retries} retries")
修改后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.service import Service from selenium.common.exceptions import TimeoutException def retry_find_element(driver, by, value, max_retries=3, timeout=15): for _ in range(max_retries): try: return WebDriverWait(driver, timeout).until(EC.element_to_be_clickable((by, value))) except TimeoutException: driver.refresh() raise TimeoutException(f"Failed to locate element after {max_retries} retries") chrome_options = Options() chrome_options.add_argument("--window-size=1920,1080") chrome_options.add_argument("--headless") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") s = Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=s, options=chrome_options) try: URL = "https://www.londonstockexchange.com/news?tab=news-explorer&period=lastmonth&page=1" driver.get(URL) # 处理Cookie弹窗 retry_find_element(driver, By.XPATH, "//button[@id='ccc-notify-accept']").click() while True: driver.get(URL) # 点击下拉框表单 dropdown_form = retry_find_element(driver, By.XPATH, "//form[contains(@class, 'js-show-results') and @action='#']") dropdown_form.click() # 选择Show 500 news选项 show_500 = retry_find_element(driver, By.XPATH, "//div[contains(@class, 'dropdown-menu')]//span[text()='Show 500 news']") show_500.find_element(By.XPATH, "..").click() # 等待表格加载完成 retry_find_element(driver, By.TAG_NAME, "table") print("Success") except Exception as e: print(f"Failure: {str(e)}") raise finally: driver.quit()
内容的提问来源于stack exchange,提问作者Aleister Crowley
相关产品推荐
相关产品推荐

