Selenium爬取SwissLipids遇StaleElementReferenceException问题求助
解决Selenium爬取SwissLipids时的StaleElementReferenceException问题
问题
使用Selenium从SwissLipids数据库批量获取脂质的InChI Key和SMILES信息时,代码第一次迭代可正常输出结果,但后续循环会抛出StaleElementReferenceException错误,无法继续执行。
报错详情
StaleElementReferenceException Traceback (most recent call last) in <cell line: 35>() 35 for row in rows: 36 # Get the elements of each row ---> 37 cells = row.find_elements('tag name', 'td') 38 39 # Check if the row contains enough cells 3 frames /usr/local/lib/python3.10/dist-packages/selenium/webdriver/remote/errorhandler.py in check_response(self, response) 243 alert_text = value["alert"].get("text") 244 raise exception_class(message, screen, stacktrace, alert_text) # type: ignore[call-arg] # mypy is not smart enough here ---> 245 raise exception_class(message, screen, stacktrace) StaleElementReferenceException: Message: stale element reference: element is not attached to the page document (Session info: headless chrome=90.0.4430.212); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#stale-element-reference-exception
原始代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time def web_driver(): options = Options() options.add_argument("--verbose") options.add_argument('--no-sandbox') options.add_argument('--headless') options.add_argument('--disable-gpu') options.add_argument('--disable-dev-shm-usage') return webdriver.Chrome(options=options) # Create an instance of the WebDriver using the configured browser options driver = web_driver() # Open the website URL in the browser url = 'https://www.swisslipids.org/#/browse_tree?entity_id=SLM:000389800' driver.get(url) # Interact with the elements on the page to extract the desired data # Find the table element table_element = driver.find_element('css selector', '.table') # Iterate through the table rows, excluding the first header row rows = table_element.find_elements('tag name', 'tr')[1:] count = 0 # Counter initialized to 0 for row in rows: # Get the elements of each row cells = row.find_elements('tag name', 'td') # Check if the row contains enough cells if len(cells) >= 2: # Extract the data from each cell identifiant_lipid = cells[0].text.strip() lipid_name = cells[1].text.strip() # Generate the specific link for each lipid ID by combining it with the base URL lipid_link = f'https://www.swisslipids.org/#/entity/{identifiant_lipid}/' count += 1 # Increment the counter print("ID lipid:", identifiant_lipid) print("Nom lipide:", lipid_name) print("Lien:", lipid_link) # Open the specific link for each lipid ID driver.get(lipid_link) time.sleep(1) # Pause for page load try: # Explicitly wait for the chemInfo element to be clickable chem_info_element = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, '#chemInfo > fieldset:nth-child(1) > div:nth-child(3)')) ) # Find all the dt and dd elements within the chemInfo element dt_elements = chem_info_element.find_elements('tag name', 'dt') dd_elements = chem_info_element.find_elements('tag name', 'dd') # Extract the desired information based on the index of the dt element for index, dt_element in enumerate(dt_elements): dt_text = dt_element.text.strip() dd_element = dd_elements[index] if dt_text == 'InChI key': inchi_key = dd_element.text.strip().replace('InChIKey=', '') print("InChI Key:", inchi_key) if dt_text == 'SMILES': smiles = dd_element.text.strip() print("SMILES:", smiles) except (NoSuchElementException, StaleElementReferenceException) as e: print("Informations chimiques non trouvées") print("Exception:", e) print() # Empty line for better readability # Print the total number of lipid IDs print("Nombre total d'ID lipid:", count) # Close the browser driver.quit()
第一次迭代正确输出示例
ID lipid: SLM:000000510 Nom lipide: hexadecanoate Lien: https://www.swisslipids.org/#/entity/SLM:000000510/ InChI Key: IPCSVZSSVZVIGE-UHFFFAOYSA-M SMILES: CCCCCCCCCCCCCCCC([O-])=O
解决方案
问题根源
当执行driver.get(lipid_link)跳转到新页面后,初始页面的DOM结构被完全替换,之前获取的rows列表中的元素引用已失效,后续循环访问这些元素时就会抛出StaleElementReferenceException。
修改方案
先在初始页面提取所有脂质的ID、名称和链接,存储到本地列表中,再遍历这个列表访问每个脂质页面,避免依赖已失效的DOM元素引用。
修改后的代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time def web_driver(): options = Options() options.add_argument("--verbose") options.add_argument('--no-sandbox') options.add_argument('--headless') options.add_argument('--disable-gpu') options.add_argument('--disable-dev-shm-usage') return webdriver.Chrome(options=options) driver = web_driver() url = 'https://www.swisslipids.org/#/browse_tree?entity_id=SLM:000389800' driver.get(url) # 先提取所有脂质信息到本地列表,避免DOM失效问题 lipids = [] # 等待表格加载完成 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CSS_SELECTOR, '.table'))) table_element = driver.find_element('css selector', '.table') rows = table_element.find_elements('tag name', 'tr')[1:] for row in rows: cells = row.find_elements('tag name', 'td') if len(cells) >= 2: identifiant_lipid = cells[0].text.strip() lipid_name = cells[1].text.strip() lipid_link = f'https://www.swisslipids.org/#/entity/{identifiant_lipid}/' lipids.append({ 'id': identifiant_lipid, 'name': lipid_name, 'link': lipid_link }) count = 0 # 遍历本地列表访问每个脂质页面 for lipid in lipids: count += 1 print("ID lipid:", lipid['id']) print("Nom lipide:", lipid['name']) print("Lien:", lipid['link']) driver.get(lipid['link']) # 等待页面加载完成,替换time.sleep更可靠 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.ID, 'chemInfo'))) try: chem_info_element = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CSS_SELECTOR, '#chemInfo > fieldset:nth-child(1) > div:nth-child(3)')) ) dt_elements = chem_info_element.find_elements('tag name', 'dt') dd_elements = chem_info_element.find_elements('tag name', 'dd') for index, dt_element in enumerate(dt_elements): dt_text = dt_element.text.strip() dd_element = dd_elements[index] if dt_text == 'InChI key': inchi_key = dd_element.text.strip().replace('InChIKey=', '') print("InChI Key:", inchi_key) if dt_text == 'SMILES': smiles = dd_element.text.strip() print("SMILES:", smiles) except (NoSuchElementException, StaleElementReferenceException) as e: print("Informations chimiques non trouvées") print("Exception:", e) print() print("Nombre total d'ID lipid:", count) driver.quit()
关键改动点
- 预提取数据到本地列表:在初始页面一次性提取所有脂质的ID、名称和链接,存储到
lipids列表中,后续循环不再依赖原始DOM元素。 - 优化等待逻辑:用
WebDriverWait等待元素出现/可见,替代固定时长的time.sleep,提升代码稳定性和效率。 - 调整元素等待条件:将
element_to_be_clickable改为visibility_of_element_located,更符合实际需求(只需元素可见即可提取文本,无需可点击)。
内容的提问来源于stack exchange,提问作者Bio-info
相关产品推荐
相关产品推荐

