使用Selenium无法抓取复杂布局表格的三个字段问题排查
Fixing Selenium Scraper for Franking Credit, Gross Dividend & Further Information Fields
Let's break down why your script is throwing that IndexError and how to fix it step by step:
Key Issues in Your Original Script
- Incorrect Button Selector: Your loop targets
td.sorting_1which is the first column of each row, not the yellow plus button that expands the hidden fields. That's why even though you're clicking elements, the hidden columns might not be consistently loaded (or loaded at all for some rows). - No Wait for Element Loading: After clicking the expand button, the browser needs a moment to render the additional columns. Your script immediately tries to access those fields before they exist, leading to shorter
tdlists and index out of range errors. - Hardcoded TD Indices: Relying on fixed indices (5,6,7) is fragile—if the table structure changes slightly, this breaks. It's better to add safeguards for row length before accessing specific positions.
Corrected Script
Here's the revised script that addresses all these issues:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = "https://www.sharedividends.com.au/mlt-dividend-history/" driver = webdriver.Chrome() driver.get(url) # Wait for the table to fully load first wait = WebDriverWait(driver, 10) table = wait.until(EC.presence_of_element_located((By.ID, "divTable"))) driver.execute_script("arguments[0].scrollIntoView();", table) # Locate and click all expand buttons (yellow plus circles) expand_buttons = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#divTable tbody tr td button.fa.fa-plus-circle"))) for button in expand_buttons: # Scroll to the button to avoid click intercept issues driver.execute_script("arguments[0].scrollIntoView();", button) # Only click if the button is still in "plus" state (not already expanded) if "fa-plus-circle" in button.get_attribute("class"): button.click() # Give the browser time to render the expanded columns time.sleep(0.5) # Now scrape the expanded rows rows = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#divTable tbody tr"))) for row in rows: tds = row.find_elements(By.TAG_NAME, "td") # Safeguard: only process rows that have all required columns if len(tds) >= 8: franking_credit = tds[5].text.strip() gross_dividend = tds[6].text.strip() further_info = tds[7].text.strip() print(f"Franking Credit: {franking_credit}, Gross Dividend: {gross_dividend}, Further Info: {further_info}") else: print("Skipping incomplete row - missing expanded columns") driver.quit()
What Changed & Why
- Accurate Button Targeting: We use
#divTable tbody tr td button.fa.fa-plus-circleto directly select the yellow plus buttons, ensuring we're clicking the right element to reveal hidden fields. - Wait Conditions:
WebDriverWaitensures the table and buttons are fully loaded before interacting with them, reducing race conditions where elements might not exist yet. - Row Length Check: Adding
if len(tds) >=8prevents theIndexErrorby skipping rows that didn't expand properly. - Small Delay After Click: A short
time.sleep(0.5)gives the browser time to render the expanded columns before we try to scrape them.
Optional Improvements
- Replace
time.sleepwith an explicit wait for the 8th TD element in the row for more reliability (avoids arbitrary delays). - Map column indices to header text dynamically (e.g., find the index of "Franking Credit" in the header row) to avoid hardcoding indices entirely.
内容的提问来源于stack exchange,提问作者MITHU
相关产品推荐
相关产品推荐

