You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML标签自身文本而非子标签内容?

解决BeautifulSoup提取标签自身文本(不含子标签)的问题

你需要提取HTML标签自身包含的直接文本,而非递归获取所有子标签的文本,避免出现内容重复。比如<div>标签只返回"I'm a div",<p>标签单独返回"I'm a paragraph"。

解决方案

核心思路是只提取标签的直接子节点中的文本节点,跳过所有子标签。可以写一个辅助函数实现这个逻辑,替换原来的tag.get_text()方法。

具体代码实现

首先导入必要模块:

from bs4 import BeautifulSoup, NavigableString

定义获取直接文本的辅助函数:

def get_direct_text(tag, strip=True):
    direct_texts = []
    # 遍历标签的直接子节点
    for child in tag.children:
        # 仅收集纯文本节点,排除子标签
        if isinstance(child, NavigableString):
            if strip:
                # 清理首尾空白,过滤空文本内容
                cleaned_text = child.strip()
                if cleaned_text:
                    direct_texts.append(cleaned_text)
            else:
                direct_texts.append(child)
    # strip模式下用空格连接有效文本,否则直接拼接原文本
    return ' '.join(direct_texts) if strip else ''.join(direct_texts)

修改原有业务代码:

html_description = """
<div>
  I'm a div
  <p>I'm a paragraph</p>
</div>
"""

soup = BeautifulSoup(html_description, 'html.parser')
TAGS_TO_APPEND = ['div', 'p', 'h1']
sanitised_description = ""

for tag in soup.find_all(True):
    if tag.name in TAGS_TO_APPEND:
        direct_text = get_direct_text(tag)
        if direct_text:  # 仅添加非空内容,避免多余换行
            sanitised_description += direct_text + '\n\n'
    elif tag.name == 'li':
        direct_text = get_direct_text(tag)
        if direct_text:
            sanitised_description += '\n* ' + direct_text

print(sanitised_description)

效果验证

运行代码后输出结果:

I'm a div

I'm a paragraph

完全符合需求:<div>仅返回自身文本,<p>文本单独提取,无重复内容。

原理说明

  • tag.children仅遍历当前标签的直接子节点,不会递归深入子标签内部
  • 通过isinstance(child, NavigableString)筛选纯文本节点,自动跳过所有子标签元素
  • 可选的strip参数可灵活控制是否清理空白、过滤空文本,适配不同格式化场景

内容的提问来源于stack exchange,提问作者Mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 03:43:26