You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式匹配Employee I时,如何提取对应前序标签的姓名?

提取HTML中职位为Employee I的员工姓名方法

我有一个包含100条以上记录的HTML文件,结构如下:

<div class="cell-62 pl-1 pt-0_5">
    <h3 class="very-big-text light-text">John Smith</h3>
        <span class="light-text">Center - VAR - Employee I</span>
</div>

<div class="cell-62 pl-1 pt-0_5">
    <h3 class="very-big-text light-text">Jenna Smith</h3>
        <span class="light-text">West - VAR - Employee I</span>
</div>

<div class="cell-62 pl-1 pt-0_5">
    <h3 class="very-big-text light-text">Jordan Smith</h3>
        <span class="light-text">East - VAR - Employee II</span>
</div>

我需要提取职位为Employee I的员工姓名,但这具有挑战性。请问如何选取后续标签包含Employee I的对应标签?是否可以使用条件实现,或是有其他方法?

现有代码片段如下:

with open("file.html", 'r') as input:
html = input.read()
    print(re.search(r'\bEmployee I\b',html).group(0))

比如,如何指定读取前一个标签?

解决方案

方法一:正则表达式(仅适用于结构固定的HTML)

如果你的HTML结构完全统一(每个员工信息都包裹在div.cell-62中,h3标签紧跟在目标span标签之前),可以用正则捕获匹配Employee I的span对应的前一个h3内容:

import re

with open("file.html", 'r') as input_file:
    html = input_file.read()

# 匹配包含Employee I的span前的h3文本
pattern = r'<h3 class="very-big-text light-text">(.*?)</h3>\s*<span class="light-text">.*?Employee I</span>'
matches = re.findall(pattern, html, re.DOTALL)

for name in matches:
    print(name.strip())

方法二:HTML解析库(推荐,稳定性更高)

正则处理HTML容易因结构微小变动失效,更可靠的方式是用BeautifulSoup解析DOM结构,通过层级关系关联姓名和职位:

  1. 先安装依赖:
pip install beautifulsoup4
  1. 提取代码:
from bs4 import BeautifulSoup

with open("file.html", 'r') as input_file:
    soup = BeautifulSoup(input_file.read(), 'html.parser')

# 遍历所有员工信息容器
for employee_div in soup.find_all('div', class_='cell-62 pl-1 pt-0_5'):
    # 获取当前容器内的职位标签
    position_span = employee_div.find('span', class_='light-text')
    # 判断职位是否为Employee I
    if position_span and 'Employee I' in position_span.text:
        # 提取对应姓名
        name_h3 = employee_div.find('h3', class_='very-big-text light-text')
        if name_h3:
            print(name_h3.text.strip())

两种方法对比

  • 正则方法:实现简单,但仅适用于HTML结构完全固定的场景,一旦标签格式、换行、属性顺序变化,就可能匹配失败。
  • BeautifulSoup方法:通过DOM层级定位元素,不受格式细节影响,适配性更强,适合处理大量或可能变动的HTML内容。

内容的提问来源于stack exchange,提问作者user15109593

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:01:39