You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium和WebDriver识别表格数据并按列抓取逐个打印?

嘿,我来帮你搞定按列抓取这个div模拟表格的数据并逐个打印的需求!

从你提供的HTML片段和截图来看,这个表格是用div元素模拟的(而非原生table标签),核心结构逻辑很清晰:

  • 整个表格容器是div.flowsheet-table
  • 每一列对应一个div.flowsheet-column
  • 列内的每一行(包括表头和数据)是div.flowsheet-row

下面用Python的BeautifulSoup库来实现这个功能,步骤很直观:

按列抓取并打印表格数据的实现方案

1. 先装依赖

如果还没安装BeautifulSoup(解析HTML用)和requests(如果需要从网页拉取HTML),先跑这个命令:

pip install beautifulsoup4 requests

2. 代码实现

场景1:用本地已有的HTML片段

我补充了模拟的完整结构(你的原始片段是截断的),直接替换成你实际的完整HTML就行:

from bs4 import BeautifulSoup

# 替换成你实际的完整HTML代码
html_content = '''
<div id="ember10911" class="ember-view flowsheet-table"> 
  <div class="header flowsheet-column"> 
    <div class="flowsheet-cell flowsheet-row"> </div> 
    <div id="ember10912" class="ember-view flowsheet-row header-row"> 
      <h4 class="header4semibold " data-ember-action="10913"> 
        <div class="ellipsis">体温</div>
      </h4>
    </div>
    <div class="flowsheet-row data-row">
      <div class="flowsheet-cell">36.5℃</div>
    </div>
    <div class="flowsheet-row data-row">
      <div class="flowsheet-cell">36.7℃</div>
    </div>
  </div>
  <div class="flowsheet-column"> 
    <div class="flowsheet-cell flowsheet-row"> </div> 
    <div class="ember-view flowsheet-row header-row"> 
      <h4 class="header4semibold "> 
        <div class="ellipsis">心率</div>
      </h4>
    </div>
    <div class="flowsheet-row data-row">
      <div class="flowsheet-cell">72次/分</div>
    </div>
    <div class="flowsheet-row data-row">
      <div class="flowsheet-cell">75次/分</div>
    </div>
  </div>
</div>
'''

# 解析HTML
soup = BeautifulSoup(html_content, 'html.parser')

# 抓所有列
columns = soup.find_all('div', class_='flowsheet-column')

# 按列遍历打印
for col_idx, column in enumerate(columns, 1):
    print(f"=== 第{col_idx}列 ===")
    # 抓当前列的所有行
    rows = column.find_all('div', class_='flowsheet-row')
    for row in rows:
        # 提取文本并去掉多余空格
        row_text = row.get_text(strip=True)
        if row_text:  # 跳过空行
            print(row_text)

场景2:直接从网页抓取表格

如果表格在线上网页里,先拉取页面内容再解析:

import requests
from bs4 import BeautifulSoup

# 替换成表格所在的实际网页URL
target_url = "你的网页地址"
# 拉取页面内容
response = requests.get(target_url)
response.raise_for_status()  # 确保请求成功

# 后续逻辑和场景1一致
soup = BeautifulSoup(response.text, 'html.parser')
columns = soup.find_all('div', class_='flowsheet-column')

for col_idx, column in enumerate(columns, 1):
    print(f"=== 第{col_idx}列 ===")
    rows = column.find_all('div', class_='flowsheet-row')
    for row in rows:
        row_text = row.get_text(strip=True)
        if row_text:
            print(row_text)

小提示

如果表头和数据行有专属class(比如header-row和data-row),可以针对性筛选,比如只想打印数据行就把rows = column.find_all('div', class_='flowsheet-row')改成rows = column.find_all('div', class_='data-row')。

内容的提问来源于stack exchange,提问作者CPP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:11:13