You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup中结合顶级无标签文本与常规标签查找功能?

解决BeautifulSoup同时提取指定标签与顶级无标签文本的问题

要实现你要的效果,直接遍历目标容器的直接子节点,分别处理标签和文本即可,代码示例如下:

from bs4 import BeautifulSoup, NavigableString, Tag

html = """
<span>
    <strong>Description</strong>
    Section1
    <ul>
        <li>line1</li>
        <li>line2</li>
        <li>line3</li>
    </ul>
    <strong>Section2</strong>
    Content2    
</span>
"""

soup = BeautifulSoup(html, 'html.parser')
target_span = soup.find('span')

result_list = []
for node in target_span.contents:
    # 处理标签节点:只保留非ul/li的标签
    if isinstance(node, Tag):
        if node.name not in ['ul', 'li']:
            result_list.append(str(node))
    # 处理文本节点:清理空白后,非空文本才保留
    elif isinstance(node, NavigableString):
        cleaned_text = node.strip()
        if cleaned_text:
            result_list.append(cleaned_text)

# 输出结果
print(result_list)

执行后会得到你期望的输出:

['<strong>Description</strong>', 'Section1', '<strong>Section2</strong>', 'Content2']

为什么之前的方法不生效?

  • select(':not(ul,li)')只会选中符合条件的标签节点,完全忽略文本节点;
  • find_all(['strong'])仅能捕获指定的strong标签,同样漏掉了顶级无标签文本。

直接遍历容器的子节点,可以同时覆盖标签和文本两种类型的节点,再通过简单过滤就能精准拿到你要的内容。

内容的提问来源于stack exchange,提问作者slothish1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 12:42:11