如何用Python/Selenium点击表格行抓取ILThermo数据库数据?
解决ILThermo左侧面板行点击问题的方法
1. 按元素索引遍历点击
左侧结果行的随机ID不影响定位,我们可以先获取所有行的元素集合,再通过索引依次点击。优先用显式等待替代硬等待,提升脚本稳定性:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 初始化等待器,最长等待10秒 wait = WebDriverWait(driver, 10) # 等待左侧结果面板加载完成,获取所有行元素 result_rows = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#dsgrid table tr"))) # 遍历每一行 for index in range(len(result_rows)): # 重新定位当前行,避免元素过期 current_row = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, f"#dsgrid table tr:nth-child({index+1})"))) current_row.click() time.sleep(1) # 等待右侧表格渲染 # 用Beautiful Soup提取右侧表格数据 soup = BeautifulSoup(driver.page_source, "html.parser") # 根据实际页面结构定位表格,示例用ID定位 target_table = soup.find("table", id="data-table") # 处理表格数据(比如提取行和列) if target_table: rows = target_table.find_all("tr") for row in rows: cols = row.find_all("td") data = [col.text.strip() for col in cols] print(data)
2. 利用共同class属性定位
如果左侧行有统一的class(比如dsgrid-row),可以直接通过class获取所有行:
# 定位所有带共同class的结果行 result_rows = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "dsgrid-row"))) for row in result_rows: row.click() time.sleep(1) # 数据提取逻辑同上...
3. 分页循环处理
要遍历所有分页的行,只需在处理完当前页后,定位下一页按钮并点击,直到没有下一页为止:
while True: # 处理当前页所有行 result_rows = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "#dsgrid table tr"))) for row in result_rows: row.click() time.sleep(1) # 数据提取... # 尝试点击下一页 try: next_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".paginator-next"))) next_btn.click() time.sleep(2) # 等待分页加载 except: # 找不到下一页按钮时退出循环 break
关键提示
- 不要过度依赖
time.sleep(),用WebDriverWait等待元素出现/可点击,能避免因页面加载慢导致的报错 - 每次循环点击前重新定位元素,可解决
StaleElementReferenceException(元素过期)问题 - 实际使用时要根据页面真实HTML结构调整CSS选择器(比如检查左侧行的父容器、右侧表格的ID或class)
内容的提问来源于stack exchange,提问作者Diego
相关产品推荐
相关产品推荐

