You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup提取指定标签文本并添加竖线分隔符

问题:如何在提取的h4与p文本间添加竖线分隔符?

现有HTML结构:

<div class="leader-info"><h4>Director of IR</h4><p>Diane PHILIPS</p></div>,
<div class="leader-info"><h4>Director of Finance</h4><p>Nancy LOPEZ</p></div>,
 <div class="leader-info"><h4>Director of HR</h4><p>George SANTOZ</p></div>

使用以下代码提取文本:

for leader_list in soup.findAll(attrs={'class':'leader-info'}):
print(leader_list.get_text())

得到的结果为:

Director of IRDiane PHILIPS
Director of FinanceNancy LOPEZ
Director of HRGeorge SANTOZ

需求是在h4和p标签的文本之间添加竖线分隔符,得到如下格式:

Director of IR|Diane PHILIPS
Director of Finance|Nancy LOPEZ
Director of HR|George SANTOZ

解决方法

方法一:分别提取h4和p标签的文本后拼接

直接定位每个leader-info下的h4和p标签,提取各自文本再用竖线连接:

for leader_list in soup.findAll(attrs={'class':'leader-info'}):
    # 提取h4文本并去除首尾空白
    position = leader_list.find('h4').get_text(strip=True)
    # 提取p文本并去除首尾空白
    name = leader_list.find('p').get_text(strip=True)
    # 拼接并打印
    print(f"{position}|{name}")

方法二:利用stripped_strings批量获取文本片段

BeautifulSoup的stripped_strings属性会返回当前标签下所有非空白的文本片段,刚好每个leader-info里有两个片段,直接用竖线拼接即可:

for leader_list in soup.findAll(attrs={'class':'leader-info'}):
    # 将文本片段转为列表后用竖线连接
    print('|'.join(list(leader_list.stripped_strings)))

两种方法都能得到需求中的格式,方法二更灵活,就算后续标签结构有小调整(比如增加同级文本标签),只要文本片段数量对应,就能直接复用。


内容的提问来源于stack exchange,提问作者John Al

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:35:04