如何通过Selenium提取指定表单pepList数据至Pandas DataFrame
Got it, let's work through this together—you've already nailed the navigation part, so now we just need to wrangle that pepList table into a Pandas DataFrame. The table does have a slightly quirky structure, but we can extract its data cleanly with targeted steps:
Step 1: Wait for the table to load (don't skip this!)
First, we need to make sure the table fully renders before trying to pull data. Using WebDriverWait avoids frustrating "element not found" errors:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for the pepList table to appear wait = WebDriverWait(driver, 10) table = wait.until(EC.presence_of_element_located((By.NAME, "pepList")))
Step 2: Grab the table headers
Let's pull the column names from the table's header section:
headers = [] # Get the header row(s) header_rows = table.find_elements(By.TAG_NAME, "thead")[0].find_elements(By.TAG_NAME, "tr") for header_row in header_rows: th_elements = header_row.find_elements(By.TAG_NAME, "th") for th in th_elements: # Clean up whitespace in header text headers.append(th.text.strip())
Step 3: Extract row data from the table body
Next, we'll loop through each row in the table body, pulling text from every cell. We'll skip empty separator rows to keep our data clean:
table_data = [] tbody_rows = table.find_elements(By.TAG_NAME, "tbody")[0].find_elements(By.TAG_NAME, "tr") for row in tbody_rows: row_data = [] td_elements = row.find_elements(By.TAG_NAME, "td") for td in td_elements: # Clean up line breaks and extra spaces in cell text cell_text = td.text.strip().replace("\n", " ") row_data.append(cell_text) # Only add rows that actually contain data if any(row_data): table_data.append(row_data)
Step 4: Convert to a Pandas DataFrame
Now we just need to combine our headers and row data into a DataFrame:
df = pd.DataFrame(table_data, columns=headers) print("Successfully extracted table! Here's a preview:") print(df.head())
Full Integrated Script
Here's how all this fits into your existing code:
from selenium import webdriver import os from selenium.webdriver.support.ui import Select, WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import pandas as pd import sys from selenium.webdriver.support import expected_conditions as EC options = Options() options.binary_location=r'C:\Program Files (x86)\Google\Chrome\Application\chrome.exe' options.add_experimental_option('excludeSwitches', ['enable-logging']) #options.add_argument("--headless") driver = webdriver.Chrome(options=options, executable_path='/mnt/c/Users/kela/Desktop/selenium/chromedriver.exe') try: driver.get('http://imed.med.ucm.es/epimhc/') # Select all desired fields (cleaned up with a loop) fields_to_select = ["mhc", "seq", "mhc_source", "class", "length", "peptide_source", "bind_level", "epitope", "epitope_level", "reference", "protein_name", "protein_source"] for field in fields_to_select: driver.find_element(By.CSS_SELECTOR, f'[value="{field}"]').click() # Trigger the search driver.find_element(By.CSS_SELECTOR, '[value=Search]').click() # Wait for table and extract data wait = WebDriverWait(driver, 10) table = wait.until(EC.presence_of_element_located((By.NAME, "pepList"))) # Extract headers headers = [] header_rows = table.find_elements(By.TAG_NAME, "thead")[0].find_elements(By.TAG_NAME, "tr") for header_row in header_rows: th_elements = header_row.find_elements(By.TAG_NAME, "th") for th in th_elements: headers.append(th.text.strip()) # Extract rows table_data = [] tbody_rows = table.find_elements(By.TAG_NAME, "tbody")[0].find_elements(By.TAG_NAME, "tr") for row in tbody_rows: row_data = [] td_elements = row.find_elements(By.TAG_NAME, "td") for td in td_elements: cell_text = td.text.strip().replace("\n", " ") row_data.append(cell_text) if any(row_data): table_data.append(row_data) # Build DataFrame df = pd.DataFrame(table_data, columns=headers) print(df.head()) finally: # Always close the driver when done driver.quit()
Quick Notes
- Waiting is critical: The table takes a moment to load after clicking Search, so
WebDriverWaitensures we don't try to extract data too early. - Cleaning text: The
strip()andreplace("\n", " ")fixes messy formatting in table cells. - Skipping empty rows: Some rows in the table are just visual separators—we skip those to keep our DataFrame tidy.
内容的提问来源于stack exchange,提问作者Slowat_Kela

