使用正则表达式匹配Employee I时,如何提取对应前序标签的姓名?
提取HTML中职位为Employee I的员工姓名方法
我有一个包含100条以上记录的HTML文件,结构如下:
<div class="cell-62 pl-1 pt-0_5"> <h3 class="very-big-text light-text">John Smith</h3> <span class="light-text">Center - VAR - Employee I</span> </div> <div class="cell-62 pl-1 pt-0_5"> <h3 class="very-big-text light-text">Jenna Smith</h3> <span class="light-text">West - VAR - Employee I</span> </div> <div class="cell-62 pl-1 pt-0_5"> <h3 class="very-big-text light-text">Jordan Smith</h3> <span class="light-text">East - VAR - Employee II</span> </div>我需要提取职位为Employee I的员工姓名,但这具有挑战性。请问如何选取后续标签包含Employee I的对应标签?是否可以使用条件实现,或是有其他方法?
现有代码片段如下:
with open("file.html", 'r') as input: html = input.read() print(re.search(r'\bEmployee I\b',html).group(0))比如,如何指定读取前一个标签?
解决方案
方法一:正则表达式(仅适用于结构固定的HTML)
如果你的HTML结构完全统一(每个员工信息都包裹在div.cell-62中,h3标签紧跟在目标span标签之前),可以用正则捕获匹配Employee I的span对应的前一个h3内容:
import re with open("file.html", 'r') as input_file: html = input_file.read() # 匹配包含Employee I的span前的h3文本 pattern = r'<h3 class="very-big-text light-text">(.*?)</h3>\s*<span class="light-text">.*?Employee I</span>' matches = re.findall(pattern, html, re.DOTALL) for name in matches: print(name.strip())
方法二:HTML解析库(推荐,稳定性更高)
正则处理HTML容易因结构微小变动失效,更可靠的方式是用BeautifulSoup解析DOM结构,通过层级关系关联姓名和职位:
- 先安装依赖:
pip install beautifulsoup4
- 提取代码:
from bs4 import BeautifulSoup with open("file.html", 'r') as input_file: soup = BeautifulSoup(input_file.read(), 'html.parser') # 遍历所有员工信息容器 for employee_div in soup.find_all('div', class_='cell-62 pl-1 pt-0_5'): # 获取当前容器内的职位标签 position_span = employee_div.find('span', class_='light-text') # 判断职位是否为Employee I if position_span and 'Employee I' in position_span.text: # 提取对应姓名 name_h3 = employee_div.find('h3', class_='very-big-text light-text') if name_h3: print(name_h3.text.strip())
两种方法对比
- 正则方法:实现简单,但仅适用于HTML结构完全固定的场景,一旦标签格式、换行、属性顺序变化,就可能匹配失败。
- BeautifulSoup方法:通过DOM层级定位元素,不受格式细节影响,适配性更强,适合处理大量或可能变动的HTML内容。
内容的提问来源于stack exchange,提问作者user15109593
相关产品推荐
相关产品推荐

