You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Selenium提取指定表单pepList数据至Pandas DataFrame

Extracting epimhc's pepList Table into a Pandas DataFrame

Got it, let's work through this together—you've already nailed the navigation part, so now we just need to wrangle that pepList table into a Pandas DataFrame. The table does have a slightly quirky structure, but we can extract its data cleanly with targeted steps:

Step 1: Wait for the table to load (don't skip this!)

First, we need to make sure the table fully renders before trying to pull data. Using WebDriverWait avoids frustrating "element not found" errors:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Wait up to 10 seconds for the pepList table to appear
wait = WebDriverWait(driver, 10)
table = wait.until(EC.presence_of_element_located((By.NAME, "pepList")))

Step 2: Grab the table headers

Let's pull the column names from the table's header section:

headers = []
# Get the header row(s)
header_rows = table.find_elements(By.TAG_NAME, "thead")[0].find_elements(By.TAG_NAME, "tr")
for header_row in header_rows:
    th_elements = header_row.find_elements(By.TAG_NAME, "th")
    for th in th_elements:
        # Clean up whitespace in header text
        headers.append(th.text.strip())

Step 3: Extract row data from the table body

Next, we'll loop through each row in the table body, pulling text from every cell. We'll skip empty separator rows to keep our data clean:

table_data = []
tbody_rows = table.find_elements(By.TAG_NAME, "tbody")[0].find_elements(By.TAG_NAME, "tr")

for row in tbody_rows:
    row_data = []
    td_elements = row.find_elements(By.TAG_NAME, "td")
    for td in td_elements:
        # Clean up line breaks and extra spaces in cell text
        cell_text = td.text.strip().replace("\n", " ")
        row_data.append(cell_text)
    # Only add rows that actually contain data
    if any(row_data):
        table_data.append(row_data)

Step 4: Convert to a Pandas DataFrame

Now we just need to combine our headers and row data into a DataFrame:

df = pd.DataFrame(table_data, columns=headers)
print("Successfully extracted table! Here's a preview:")
print(df.head())

Full Integrated Script

Here's how all this fits into your existing code:

from selenium import webdriver
import os
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import pandas as pd
import sys
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.binary_location=r'C:\Program Files (x86)\Google\Chrome\Application\chrome.exe'
options.add_experimental_option('excludeSwitches', ['enable-logging'])
#options.add_argument("--headless")
driver = webdriver.Chrome(options=options, executable_path='/mnt/c/Users/kela/Desktop/selenium/chromedriver.exe')

try:
    driver.get('http://imed.med.ucm.es/epimhc/')
    
    # Select all desired fields (cleaned up with a loop)
    fields_to_select = ["mhc", "seq", "mhc_source", "class", "length", "peptide_source", 
                       "bind_level", "epitope", "epitope_level", "reference", "protein_name", "protein_source"]
    for field in fields_to_select:
        driver.find_element(By.CSS_SELECTOR, f'[value="{field}"]').click()
    
    # Trigger the search
    driver.find_element(By.CSS_SELECTOR, '[value=Search]').click()
    
    # Wait for table and extract data
    wait = WebDriverWait(driver, 10)
    table = wait.until(EC.presence_of_element_located((By.NAME, "pepList")))
    
    # Extract headers
    headers = []
    header_rows = table.find_elements(By.TAG_NAME, "thead")[0].find_elements(By.TAG_NAME, "tr")
    for header_row in header_rows:
        th_elements = header_row.find_elements(By.TAG_NAME, "th")
        for th in th_elements:
            headers.append(th.text.strip())
    
    # Extract rows
    table_data = []
    tbody_rows = table.find_elements(By.TAG_NAME, "tbody")[0].find_elements(By.TAG_NAME, "tr")
    for row in tbody_rows:
        row_data = []
        td_elements = row.find_elements(By.TAG_NAME, "td")
        for td in td_elements:
            cell_text = td.text.strip().replace("\n", " ")
            row_data.append(cell_text)
        if any(row_data):
            table_data.append(row_data)
    
    # Build DataFrame
    df = pd.DataFrame(table_data, columns=headers)
    print(df.head())
    
finally:
    # Always close the driver when done
    driver.quit()

Quick Notes

  • Waiting is critical: The table takes a moment to load after clicking Search, so WebDriverWait ensures we don't try to extract data too early.
  • Cleaning text: The strip() and replace("\n", " ") fixes messy formatting in table cells.
  • Skipping empty rows: Some rows in the table are just visual separators—we skip those to keep our DataFrame tidy.

内容的提问来源于stack exchange,提问作者Slowat_Kela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:37:12