为何Chrome二次爬取表格列文本时触发StaleElementReferenceException?
Hey there, let's tackle your Chrome-specific scraping issue and optimize that table extraction at the same time.
First, let's break down why that second scrape fails on Chrome 74: it's likely a combination of leftover frame context and how the old chromedriver caches element text after the first successful scrape. Your workaround of switching pages works because it forces a full DOM refresh in the target frame, but we can fix this without manual page hops. Plus, your current nested-loop approach for extracting table data is inefficient—let's fix both problems.
Fixes for Chrome's Second-Scrape Failure
- Reset Frame Context Every Time: Always switch back to the default content before navigating to frames again. This avoids lingering context from the first scrape that might confuse the driver.
- Use Stable Text Extraction: Replace
col.textwithtextContent(via JavaScript) instead—sometimestextrelies on element visibility, whiletextContentpulls directly from the DOM, which is more reliable for cached elements. - Wait for Fully Loaded Elements: Use explicit waits for clickable/visible elements instead of immediate
find_elementcalls, ensuring the DOM is ready before interacting.
Optimized Table Scraping Method
Instead of looping through each row and column with WebDriver calls (which is slow due to constant browser-driver communication), we can extract the entire table in one go using JavaScript. This cuts down on round trips and makes the code cleaner.
Here's the revised function incorporating all these fixes:
import time from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By def Get_Faults_List(Port_Number=None, PSU=None, Retries=5): for attempt in range(Retries): try: # Start fresh by resetting to default content every attempt self.driver.switch_to.default_content() # Navigate to the relevant fault view if Port_Number: Device_Panel_Frame.Click_Port(self, Port_Number) elif PSU: if not Device_Panel_Frame.Click_PSU(self, PSU): return None Left_Panel_Frame.Click_Fault(self) # Re-enter frames with explicit waits to ensure load completion self.driver.switch_to.default_content() main_body = WebDriverWait(self.driver, 3).until( EC.presence_of_element_located((By.NAME, 'main_page')) ) self.driver.switch_to.frame(main_body) # Wait for alarms tab to be clickable before interacting alarms_tab = WebDriverWait(self.driver, 3).until( EC.element_to_be_clickable((By.ID, 'tab_alarms')) ) alarms_tab.click() # Switch to faults frame only after it's fully present fault_screen = WebDriverWait(self.driver, 3).until( EC.presence_of_element_located((By.NAME, 'faults')) ) self.driver.switch_to.frame(fault_screen) # Optional: Refresh the frame to clear Chrome's cache (for persistent issues) # self.driver.execute_script("location.reload();") # WebDriverWait(self.driver, 3).until( # EC.presence_of_element_located((By.ID, 'faultTab')) # ) # Extract entire table data in one JS call (way faster!) current_faults = self.driver.execute_script(""" const table = document.getElementById('faultTab'); // Target only relevant rows (exclude those with id/style attributes) const rows = table.querySelectorAll('tr:not([id]):not([style])'); const tableData = []; rows.forEach(row => { const columns = row.querySelectorAll('td'); // Convert TD elements to trimmed text values const rowData = Array.from(columns).map(col => col.textContent.trim()); tableData.push(rowData); }); return tableData; """) # Return valid data immediately if found if current_faults: return current_faults raise ValueError("No valid fault rows extracted") except Exception as e: print(f"Attempt {attempt + 1} failed: {str(e)}. {Retries - (attempt + 1)} retries left.") self.Refresh_Screen() time.sleep(5) continue print("All retries exhausted—failed to extract fault list") return None
Extra Chrome-Specific Fix (If Needed)
If the issue persists even after the above changes, disable Chrome's cache when initializing the driver to prevent stale DOM data from being reused:
from selenium import webdriver chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--disable-cache") chrome_options.add_argument("--disable-application-cache") # Initialize your driver with these options self.driver = webdriver.Chrome(options=chrome_options, executable_path="path/to/chromedriver-74")
Why This Works Better
- Faster Extraction: The JavaScript approach pulls all table data in a single browser call, eliminating the overhead of dozens of
find_elementrequests. - More Reliable: Resetting frame context and using
textContentavoids Chrome's caching quirks that caused the second scrape to fail. - Cleaner Code: Less nested logic makes it easier to debug and maintain.
内容的提问来源于stack exchange,提问作者Moshe S.

