You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的BeautifulSoup提取HTML中的指定无标签文本

提取无独立标签的Text2方案

先看目标HTML结构:

<h1 class="col">
  <small>Text1</small>
  Text2
</h1>

用BeautifulSoup提取Text2有这几种实用方法:

  • 筛选纯文本节点
    通过contents获取h1标签下的所有子节点,过滤出纯文本类型的节点,再清理多余空白:
from bs4 import BeautifulSoup, NavigableString

html = '''
<h1 class="col">
  <small>Text1</small>
  Text2
</h1>
'''

soup = BeautifulSoup(html, 'html.parser')
h1_col = soup.find('h1', class_='col')
# 挑出非标签的文本节点,去掉空内容
text2 = [node.strip() for node in h1_col.contents if isinstance(node, NavigableString) and node.strip()]
print(text2[0])  # 输出Text2
  • 排除small标签内容
    先提取h1的所有文本,再剔除small标签里的内容:
h1_col = soup.find('h1', class_='col')
small_content = h1_col.find('small').get_text(strip=True)
# 分割文本后排除small的内容
text2 = [t for t in h1_col.get_text(strip=True).split() if t != small_content][0]
print(text2)  # 输出Text2
  • 文本替换法
    直接获取h1的完整文本,把small标签的内容替换为空,剩下的就是目标Text2:
h1_full_text = h1_col.get_text(strip=True)
small_content = h1_col.find('small').get_text(strip=True)
text2 = h1_full_text.replace(small_content, '').strip()
print(text2)  # 输出Text2

内容的提问来源于stack exchange,提问作者LennardDerudder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:21:35