You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup移除所有HTML标签(除<i>斜体标签外)并保留换行信息

解决BeautifulSoup保留文本、换行和格式的问题

首先,咱们来分析你之前操作的问题所在:

你之前的错误点

  1. 使用get_text('\n')的问题
    这个方法会直接丢弃所有HTML标签信息(包括<i>的标识),而且默认会在每个标签边界添加换行,导致原本连续的文本被不合理拆分;同时它也没处理HTML实体(比如&amp;不会转换成&)和&nbsp;,所以结果既丢失了格式,又有乱码和多余空行。

  2. 用find_all('div')拆分的问题
    原始HTML存在嵌套<div>和空<div>(比如只包含<br>的div),这会导致提取的文本出现空行或重复内容;另外,这个方法无法合并跨标签的连续文本(比如开头的Tip:内容分散在3个<i>标签里,但它们属于同一逻辑行),也没处理HTML实体和特殊空格。

正确的解决方案

根据你的需求(移除所有HTML元素,保留所有文本、正确换行,合并分散的连续文本),可以按以下步骤处理:

from bs4 import BeautifulSoup
import html

data = '&lt;div&gt;&lt;div&gt;&lt;i&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;Tip: amplitude is decreased in axonal neuropathies; CV&amp;nbsp;&lt;/font&gt;&lt;/i&gt;&lt;i&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;&amp;amp; latency are&lt;/font&gt;&lt;/i&gt;&lt;i&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;&amp;nbsp;prolonged in&amp;nbsp;demyelination&lt;/font&gt;&lt;/i&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;i&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;&lt;br&gt;&lt;/font&gt;&lt;/i&gt;&lt;/div&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;&lt;i&gt;1. Onset latency:&lt;/i&gt;&amp;nbsp;&lt;/font&gt;is the time required for an electrical stimulus to initiate an evoked potential. This reflects the conduction along the fastest fibers.&amp;nbsp;&lt;div&gt;&lt;i&gt;- Prolonged in demyelination.&lt;/i&gt;&lt;br&gt;&lt;div&gt;&lt;div&gt;&lt;br&gt;&lt;/div&gt;&lt;div&gt;&lt;font color=&quot;&quot;#ff086c&quot;&quot;&gt;&lt;i&gt;2. Peak latency:&lt;/i&gt;&amp;nbsp;&lt;/font&gt;represents the latency along the majority of the axons and is measured at the peak of the waveform amplitude.&amp;nbsp;&lt;/div&gt;&lt;div&gt;&lt;i&gt;- Prolonged in demyelination.&lt;/i&gt;'

# 1. 解析转义后的HTML内容
soup = BeautifulSoup(data, 'html.parser')

# 2. 将<br>标签替换为换行符,还原原始换行逻辑
for br_tag in soup.find_all('br'):
    br_tag.replace_with('\n')

# 3. 获取文本内容,用换行分隔各个块
raw_text = soup.get_text('\n')

# 4. 处理HTML实体(比如把&amp;转换成&,修复特殊字符)
processed_text = html.unescape(raw_text)

# 5. 替换非空格转义字符为普通空格,统一空格格式
processed_text = processed_text.replace('\xa0', ' ')

# 6. 过滤空行,清理每行前后的多余空格
final_lines = [line.strip() for line in processed_text.split('\n') if line.strip()]

# 7. 输出最终整理后的文本
print('\n'.join(final_lines))

代码说明

  • 先替换<br>为换行符,确保原始的换行意图被准确保留。
  • 使用html.unescape()处理HTML实体,解决&amp;、&nbsp;这类转义字符的显示问题。
  • 最后过滤空行和多余空格,得到整洁的结构化文本。

运行这段代码后,你会得到完全符合期望的输出:

Tip: amplitude is decreased in axonal neuropathies; CV & latency are prolonged in demyelination
1. Onset latency: is the time required for an electrical stimulus to initiate an evoked potential. This reflects the conduction along the fastest fibers.
- Prolonged in demyelination.
2. Peak latency: represents the latency along the majority of the axons and is measured at the peak of the waveform amplitude.
- Prolonged in demyelination.

如果你还想保留<i>标签对应的斜体格式(用Markdown的*标记),可以调整代码为递归遍历节点,给<i>内的文本添加斜体标记,有需要的话可以随时问我!

内容的提问来源于stack exchange,提问作者Code Monkey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 20:22:36