You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium WebDriver如何遍历表格/网格隐藏行并获取值

Solution for Capturing Data from Virtual Scrolling Tables with Selenium

Got it, let's tackle this virtual scrolling table issue—those dynamically reused rows can be tricky! Here's a practical approach to capture all row data, even the ones that stay hidden until you scroll:

Core Idea

Virtualized tables only keep a small set of row elements (in your case, 20) in the DOM at any time. As you scroll, these rows get repurposed to show new data instead of creating fresh elements. So instead of trying to find all rows at once, we'll:

  1. Grab data from the currently visible rows
  2. Scroll to trigger the table to load the next batch of data (reusing existing rows)
  3. Repeat until we've captured all unique rows

Working Code Example (Python)

This uses ChromeDriver, but you can adapt it to other browsers easily:

from selenium import webdriver
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Initialize driver and navigate to your target page
driver = webdriver.Chrome()
driver.get("your_target_page_url_here")

# Wait for the table to fully load (replace selector with your table's container)
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "table"))
)

all_row_data = []
row_selector = (By.CSS_SELECTOR, "table tr")  # Adjust to match your row selector
# Use your table's dedicated scroll container if it exists (instead of body)
scroll_container = driver.find_element(By.TAG_NAME, "body")

while True:
    # Get all currently visible rows
    current_rows = WebDriverWait(driver, 5).until(
        EC.visibility_of_all_elements_located(row_selector)
    )
    
    # Extract data from each row (adjust cell selector to match your table's structure)
    current_batch = [
        [cell.text.strip() for cell in row.find_elements(By.TAG_NAME, "td")]
        for row in current_rows
    ]
    
    # Add only new data to avoid duplicates (from row reuse)
    new_entries = 0
    for row in current_batch:
        if row not in all_row_data:
            all_row_data.append(row)
            new_entries += 1
    
    # Stop looping if no new data loads (we've reached the end of the table)
    if new_entries == 0:
        break
    
    # Scroll to the last visible row to trigger new data loading
    last_row = current_rows[-1]
    ActionChains(driver).move_to_element(last_row).perform()
    
    # Wait for the table to update row data (adjust sleep time based on your page's speed)
    time.sleep(1)

# Output the results
print(f"Total rows captured: {len(all_row_data)}")
for idx, row in enumerate(all_row_data, 1):
    print(f"Row {idx}: {row}")

driver.quit()

Key Tips for Reliability

  • Use Explicit Waits: Replace time.sleep with WebDriverWait whenever possible—it waits until elements are ready instead of relying on fixed time delays, making the script more stable.
  • Tweak Selectors: Swap out the CSS selectors for table, tr, and td to match your actual page's HTML. If your table has a dedicated scroll container (like a div with overflow: auto), use that instead of the body element.
  • Alternative Scrolling: If move_to_element doesn't trigger new rows, try scrolling by pixels with JavaScript:
    driver.execute_script("arguments[0].scrollTop += 500;", scroll_container)
    
  • Deduplication is Critical: Since rows are reused, you'll often see overlapping data between batches. Checking if a row is already in all_row_data ensures you don't end up with duplicate entries.

内容的提问来源于stack exchange,提问作者fummy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:27:25