You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从HTML文件中提取所有嵌套文本(尽量不依赖库)

提取HTML中所有文本的两种方案

方案一:纯正则实现(无第三方库)

如果不想引入第三方库,可以用正则匹配并移除所有HTML标签,再清理多余空白来提取文本。核心思路是把所有<>包裹的标签内容替换为空,剩下的就是文本内容。

示例代码:

import re

html_content = """<div>
  <div>
    <h1>Title</h1>
    <p>Some text i want to fetch</p>
  </div>
  <RandomVueComponent>Some other text i want to fetch</RandomVueComponent>
</div>
<a href="#">Might be some text too</a>
"""

# 移除所有HTML标签
clean_text = re.sub(r'<[^>]+>', '', html_content)
# 清理多余的换行、空格,过滤空行
text_lines = [line.strip() for line in clean_text.split('\n') if line.strip()]
# 合并成最终文本(也可以保留列表形式按行输出)
final_text = '\n'.join(text_lines)

print(final_text)

输出结果:

Title
Some text i want to fetch
Some other text i want to fetch
Might be some text too

注意:正则处理HTML有局限性,如果HTML里存在标签属性包含>(比如<div data-x="a>b">)、注释、<script>/<style>这类不需要提取的文本块,正则会失效。这种情况下更推荐用下面的方案。

方案二:用BeautifulSoup(无需指定标签)

你之前对BeautifulSoup的理解有误——它完全不需要以特定标签作为匹配依据,直接调用get_text()就能提取所有嵌套层级的文本,不管标签类型。而且它能正确处理复杂HTML结构,避免正则的各种坑。

示例代码:

from bs4 import BeautifulSoup

html_content = """<div>
  <div>
    <h1>Title</h1>
    <p>Some text i want to fetch</p>
  </div>
  <RandomVueComponent>Some other text i want to fetch</RandomVueComponent>
</div>
<a href="#">Might be some text too</a>
"""

soup = BeautifulSoup(html_content, 'html.parser')
# 提取所有文本,参数strip=True会自动清理多余空白,separator指定分隔符
clean_text = soup.get_text(strip=True, separator='\n')

print(clean_text)

输出结果和正则方案一致,但能处理更复杂的HTML场景,比如包含注释、特殊属性的标签等。

内容的提问来源于stack exchange,提问作者Pous17

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 03:11:22