You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup处理嵌套在<p>中的<strong>标签提取标题内容

问题

现有Python代码可提取独立<strong>标签文本作为标题,后续同级元素作为对应内容,但当<strong>嵌套在<p>标签内时失效。需修改代码实现:以<strong>标签内文本为键,下一个<strong>标签之前的元素为值,适配两种场景。

输入示例1(独立<strong>)

example = """<strong>First Title</strong><p>Content of first title</p><p>Content of first title</p><strong>Second title</strong><p>Content of second title</p>"""

输入示例2(<strong>嵌套在<p>内)

example = """<p><strong>First Title</strong></p><p>Content of first title</p><p>Content of first title</p><p><strong>Second title</strong></p><p>Content of second title</p>"""

期望输出

{'First Title': '<p>Content of first title</p> <p>Content of first title</p>', 'Second title': '<p>Content of second title</p>'}
解决方案

核心思路是先统一两种场景的目标节点:不管<strong>是独立存在还是嵌套在<p>内,都将其所在的同级节点纳入处理列表,再遍历节点收集后续内容。

适配代码如下:

from bs4 import BeautifulSoup

def extract_title_content(html):
    soup = BeautifulSoup(html, 'html.parser')
    # 整理所有包含<strong>的目标节点
    target_nodes = []
    for strong_tag in soup.find_all('strong'):
        parent = strong_tag.parent
        # 若<strong>嵌套在<p>内,取<p>作为目标节点;否则直接取<strong>
        if parent.name == 'p' and parent.find('strong') == strong_tag:
            target_nodes.append(parent)
        else:
            target_nodes.append(strong_tag)
    
    result = {}
    for idx, current_node in enumerate(target_nodes):
        # 提取标题文本
        title = current_node.get_text(strip=True)
        content_elements = []
        next_sibling = current_node.next_sibling
        
        # 收集当前节点到下一个目标节点之间的所有有效内容
        while next_sibling is not None:
            # 遇到下一个目标节点则停止
            if next_sibling in target_nodes:
                break
            # 跳过空文本节点
            if isinstance(next_sibling, str) and next_sibling.strip() == '':
                next_sibling = next_sibling.next_sibling
                continue
            # 将元素转为字符串存入列表
            content_elements.append(str(next_sibling))
            next_sibling = next_sibling.next_sibling
        
        # 拼接内容字符串
        result[title] = ' '.join(content_elements)
    return result

# 测试示例1
example1 = """<strong>First Title</strong><p>Content of first title</p><p>Content of first title</p><strong>Second title</strong><p>Content of second title</p>"""
print(extract_title_content(example1))

# 测试示例2
example2 = """<p><strong>First Title</strong></p><p>Content of first title</p><p>Content of first title</p><p><strong>Second title</strong></p><p>Content of second title</p>"""
print(extract_title_content(example2))

代码说明

  • 先遍历所有<strong>标签,判断其是否嵌套在<p>内,统一两种场景的处理节点,避免因层级差异导致失效。
  • 遍历每个目标节点时,逐个检查后续兄弟节点,直到遇到下一个目标节点为止,期间跳过空文本节点,只收集有效元素。
  • 将收集到的元素转为字符串后拼接,作为对应标题的内容值。

内容的提问来源于stack exchange,提问作者Rakesh Shetty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:01:15