如何用Selenium循环点击并结合BeautifulSoup4抓取隐藏表格数据
Alright, let's tackle this problem head-on. I've worked through similar scenarios before, so here's a step-by-step solution that covers clicking expand arrows, scraping table data, handling missing values, and exporting to Excel using Selenium and BeautifulSoup.
1. Setup Dependencies
First, make sure you have all required libraries installed. Run this command in your terminal:
pip install selenium beautifulsoup4 pandas
Also, download the matching browser driver (e.g., ChromeDriver for Chrome) and ensure it's in your system PATH or specify its path in the code.
2. Loop to Click Expand Arrows
The key here is reliable element interaction—we'll use explicit waits to avoid race conditions, and handle common exceptions like stale elements or blocked clicks. Adjust the CSS selectors to match your actual HTML structure.
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import ElementClickInterceptedException, StaleElementReferenceException from selenium.webdriver.common.action_chains import ActionChains # Initialize browser and navigate to target page driver = webdriver.Chrome() driver.get("YOUR_TARGET_PAGE_URL") wait = WebDriverWait(driver, 10) # Helper to get fresh list of expand arrows def get_expand_arrows(): return wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".expand-arrow"))) # Replace with your arrow selector # Loop through each arrow and click to expand tables arrows = get_expand_arrows() for idx, arrow in enumerate(arrows): try: # Scroll to arrow to avoid overlay blocks ActionChains(driver).move_to_element(arrow).perform() # Wait for arrow to be clickable, then click wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, f".expand-arrow:nth-of-type({idx+1})"))).click() # Wait for the corresponding table to load (replace selector with your table's class) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, f".expanded-table:nth-of-type({idx+1})"))) except (ElementClickInterceptedException, StaleElementReferenceException): # Retry if element becomes stale or blocked arrows = get_expand_arrows() ActionChains(driver).move_to_element(arrows[idx]).perform() wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, f".expand-arrow:nth-of-type({idx+1})"))).click() wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, f".expanded-table:nth-of-type({idx+1})")))
3. Scrape Table Data (13 Rows per Table)
Once all tables are expanded, parse the page with BeautifulSoup to extract rows. We'll ensure we get exactly 13 rows per table—filling empty values if a table has fewer rows.
from bs4 import BeautifulSoup import pandas as pd # Parse the page source soup = BeautifulSoup(driver.page_source, "html.parser") # Store all scraped data all_table_rows = [] # Iterate over each expanded table tables = soup.select(".expanded-table") # Replace with your table selector for table in tables: # Get all rows in the current table rows = table.find_all("tr") # Extract up to 13 rows; fill empty lists for missing rows table_data = [] for row_idx in range(13): if row_idx < len(rows): # Extract text from each cell, stripping extra whitespace cells = [cell.get_text(strip=True) for cell in rows[row_idx].find_all(["td", "th"])] table_data.append(cells) else: # Use the column count from the first row to create empty cells col_count = len(table_data[0]) if table_data else 5 # Adjust default col count if needed table_data.append([""] * col_count) # Add current table's data to the main list all_table_rows.extend(table_data)
4. Export to Excel with Missing Values
Use Pandas to convert the scraped data into a DataFrame and export it to Excel—missing values will automatically be left blank.
# Define your table column names (adjust to match your actual columns) column_names = ["Column 1", "Column 2", "Column 3", "Column 4", "Column 5"] # Create DataFrame df = pd.DataFrame(all_table_rows, columns=column_names) # Export to Excel (index=False removes the auto-generated row numbers) df.to_excel("scraped_table_data.xlsx", index=False) # Clean up: close the browser driver.quit()
Key Notes for Your Use Case
- Selector Adjustments: Replace all CSS selectors (
.expand-arrow,.expanded-table) with the actual selectors from your HTML code. Use browser dev tools to inspect elements and get the right selectors. - Pagination: If you need to handle multiple pages, add a loop to click the "Next Page" button after processing each page, then repeat the arrow-click and scraping steps.
- Wait Times: Adjust the
WebDriverWaittimeout (currently 10 seconds) if your page loads slower.
内容的提问来源于stack exchange,提问作者Jamal Hajjar

