求助解决Selenium爬取3D打印机价格时的StaleElementReferenceException错误
Hey there, let's break down and fix the issue you're hitting with your web crawler. That StaleElementReferenceException is a common gotcha in Selenium—it happens when the element you've grabbed a reference to gets disconnected from the page's DOM (usually because the page refreshed, re-rendered, or the element was replaced dynamically after you fetched it).
Let's walk through the fixes step by step:
First, Fix the Import & Add Proper Waiting
Your code is missing the import for webdriver, and you're not waiting for the page to fully load before trying to grab elements. This is a big reason you might be hitting stale elements or missing the price element entirely. We'll use explicit waits to make sure Selenium waits until the price element is available (or times out if it's not there).
Revised Code with Fixes
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, StaleElementReferenceException def get_printer_price(): driver = webdriver.Chrome() try: url = 'https://www.dx.com/p/creality-cr10-v2-upgrade-ultraquiet-twoway-sphenoid-cooling-3d-printer-eu-plug-2711457.html#.Xy6c1SgzZhE' driver.get(url) # Wait up to 10 seconds for the price element to be present wait = WebDriverWait(driver, 10) # Use a loop to handle potential stale element issues for _ in range(3): # Retry up to 3 times if stale try: price_element = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'low-sale-price'))) price_text = price_element.text.strip() if price_text: print(f"Printer price: {price_text}") else: print('Printer out of stock') break except StaleElementReferenceException: continue # Retry if element went stale except TimeoutException: print('Price element not found—printer is likely out of stock or page structure changed') finally: driver.quit() # Make sure to close the browser when done get_printer_price()
Key Improvements Explained
- Explicit Waits:
WebDriverWaitensures we wait for the price element to load instead of trying to access it immediately (which fails if the page is still loading). - Stale Element Retry Loop: We wrap the element access in a loop that retries up to 3 times if we hit a stale reference—this handles cases where the page re-renders right after we try to grab the element.
- Proper Cleanup: The
finallyblock ensures the browser closes even if an error occurs, preventing leftover Chrome processes. - Clearer Logic: We check if the price text is empty to determine if the printer is out of stock, and handle cases where the element never loads (timeout).
Additional Notes
- If the page uses dynamic JavaScript to load the price (e.g., after scrolling or a delay), you might need to adjust the wait conditions—for example, using
EC.visibility_of_element_locatedinstead ofpresence_of_element_locatedto ensure the element is visible on the page. - Always double-check the page's HTML structure to make sure the
low-sale-priceclass is still the correct selector for the price element—websites often update their layouts!
内容的提问来源于stack exchange,提问作者igor

