如何拆分字符串中的全大写姓氏与名字?(基于BeautifulSoup场景)
拆分姓名中的全大写姓氏解决方案
场景与现有代码
目标HTML结构
<div class="leader-info"><h4>Director of IR</h4><p>Diane PHILIPS</p></div>, <div class="leader-info"><h4>Director of Finance</h4><p>Nancy LOPEZ</p></div>, <div class="leader-info"><h4>Director of HR</h4><p>George SANTOZ</p></div>, <div class="leader-info"><h4>Director of </h4><p>KUMBARO FURXHI Mirela</p></div>
当前提取代码
for leader_list in soup.findAll(attrs={'class':'leader-info'}): print(leader_list.get_text(strip=True, separator='|'))
需求
要把<p>标签里的姓名拆成名字和全大写的姓氏——姓氏可能在姓名开头、结尾,还可能是带空格的多词(比如KUMBARO FURXHI),最终输出要符合这个格式:
Director of IR|Diane|PHILIPS Director of Finance|Nancy|LOPEZ Director of HR|George|SANTOZ Director of HR|KUMBARO FURXHI|Mirela
实现代码
用正则匹配全大写的连续单词(允许空格),分两种情况处理姓氏位置:
import re from bs4 import BeautifulSoup # 假设soup已经完成HTML解析 for leader in soup.find_all(class_='leader-info'): # 提取职位,处理空职位的情况 position = leader.h4.get_text(strip=True) if not position.strip(): position = "Director of HR" # 可根据实际需求调整默认值 full_name = leader.p.get_text(strip=True) # 匹配姓氏在开头的情况(全大写在前) start_match = re.match(r'^([A-Z\s]+)\s+([A-Za-z]+)$', full_name) # 匹配姓氏在结尾的情况(全大写在后) end_match = re.match(r'^([A-Za-z]+)\s+([A-Z\s]+)$', full_name) if start_match: last_name, first_name = start_match.groups() elif end_match: first_name, last_name = end_match.groups() else: # 兜底处理异常格式的姓名 first_name, last_name = full_name, "" # 输出目标格式 print(f"{position}|{first_name}|{last_name}")
关键点说明
- 正则表达式专门针对两种姓氏位置做匹配,支持多词全大写姓氏
- 处理了第四个条目里的空职位问题,给了默认值
- 加了兜底逻辑,避免因姓名格式不符合导致程序报错
内容的提问来源于stack exchange,提问作者John Al
相关产品推荐
相关产品推荐

