You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:稳健移除HTML换行与缩进并按浏览器规则处理标签内内容

解决HTML标签内空白字符预处理的两种方法

一、用现成库快速实现

推荐使用htmlmin库,它专门用于压缩HTML,默认会把标签内的换行、缩进、连续空格这类空白字符替换成单个空格,同时保留HTML结构完整性。

  1. 安装库:
pip install htmlmin
  1. 使用示例:
import htmlmin

# 原始HTML内容
raw_html = """
<div>
    Hello
    World
    <p>
        This is a test
    </p>
</div>
"""

# 预处理HTML,remove_all_empty_space=False避免标签粘在一起
compressed_html = htmlmin.minify(raw_html, remove_all_empty_space=False)
print(compressed_html)

处理后输出:

<div> Hello World <p> This is a test </p> </div>

二、自行编写正则替换

不想额外装库的话,用Python内置的re模块就能实现,核心是把连续的空白字符(换行、制表符、多个空格)替换成单个空格。

示例代码:

import re

raw_html = """
<div>
    Hello
    World
    <p>
        This is a test
    </p>
</div>
"""

# 替换所有连续空白字符为单个空格,strip()去掉首尾多余空格
processed_html = re.sub(r'\s+', ' ', raw_html).strip()
print(processed_html)

如果需要保留<pre>标签内的空白格式,可以先拆分HTML单独处理这类标签,不过大部分场景下上面的简单正则足够满足需求。

另外补充:如果只是想用BeautifulSoup提取干净文本,直接优化get_text的参数就行:

from bs4 import BeautifulSoup

soup = BeautifulSoup(raw_html, 'html.parser')
clean_text = soup.get_text(strip=True, separator=' ')
print(clean_text)

这种方式是直接提取紧凑格式的文本,而非预处理HTML文件,可根据需求选择。

内容的提问来源于stack exchange,提问作者clel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 01:13:17