遍历动态可迭代对象时遇StaleElementReferenceException的优化解法问询
Hey there! I’ve dealt with this exact stale element headache when scraping dynamically loaded sites like justjoin.it, so I know how frustrating it is when your scroll stops halfway with that error. Let’s break down the better solutions beyond just wrapping things in try/except.
Why Your Current Code Breaks
Even though you know the root cause, let’s recap quickly: When you scroll on justjoin.it, the site loads new job listings dynamically. This updates the DOM, which invalidates the WebElement references you grabbed at the start of your loop. Those old elements are no longer attached to the page, hence the stale error. Try/except just skips the broken elements but doesn’t fix the fact that your entire list of elements is now outdated.
Solution 1: Re-fetch Elements Every Loop (For Targeted Scrolling)
Instead of grabbing all elements once at the top, re-fetch the list of elements inside your while loop. This ensures you’re always working with fresh, valid references to elements in the current DOM. We’ll also add a check to stop when we’ve reached the bottom of the page (no more new content to load):
from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() # Or your preferred driver driver.get('https://justjoin.it/') driver.maximize_window() last_page_height = driver.execute_script("return document.body.scrollHeight") while True: # Re-fetch the elements every time to avoid stale references job_elements = driver.find_elements(By.CLASS_NAME, 'css-1x9zltl') for element in job_elements: # Scroll to the element (smooth scroll is optional but nicer) driver.execute_script("arguments[0].scrollIntoView({behavior: 'smooth'});", element) # Add a small delay to let any lazy-loaded content render time.sleep(0.5) # Check if we've reached the bottom (no new content loaded) new_page_height = driver.execute_script("return document.body.scrollHeight") if new_page_height == last_page_height: break last_page_height = new_page_height driver.quit()
Solution 2: Scroll to Bottom Directly (For Loading All Content)
If you don’t need to scroll to each element individually and just want to load all job listings, a simpler approach is to repeatedly scroll to the bottom of the page until no new content loads. This avoids element references entirely:
from selenium import webdriver import time driver = webdriver.Chrome() driver.get('https://justjoin.it/') driver.maximize_window() last_page_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll all the way to the bottom driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new content to load (adjust the sleep time if needed) time.sleep(2) new_page_height = driver.execute_script("return document.body.scrollHeight") if new_page_height == last_page_height: break last_page_height = new_page_height driver.quit()
Pro Tip: Replace time.sleep with Explicit Waits
Instead of using fixed time.sleep delays (which can be unreliable), use Selenium’s WebDriverWait to wait for new elements to load. This makes your code more robust:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Inside your loop, after scrolling: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'css-1x9zltl')) )
This waits up to 10 seconds for a new job element to appear before proceeding, so you don’t waste time waiting when content loads quickly.
Key Takeaway
The main issue with your original code is reusing stale WebElement references. By re-fetching elements each loop or switching to a bottom-scroll approach, you eliminate the stale element problem entirely. Explicit waits add an extra layer of reliability for dynamic content.
内容的提问来源于stack exchange,提问作者beginsql

