如何用Beautiful Soup处理嵌套在<p>中的<strong>标签提取标题内容
问题
现有Python代码可提取独立<strong>标签文本作为标题,后续同级元素作为对应内容,但当<strong>嵌套在<p>标签内时失效。需修改代码实现:以<strong>标签内文本为键,下一个<strong>标签之前的元素为值,适配两种场景。
输入示例1(独立<strong>)
example = """<strong>First Title</strong><p>Content of first title</p><p>Content of first title</p><strong>Second title</strong><p>Content of second title</p>"""
输入示例2(<strong>嵌套在<p>内)
example = """<p><strong>First Title</strong></p><p>Content of first title</p><p>Content of first title</p><p><strong>Second title</strong></p><p>Content of second title</p>"""
期望输出
{'First Title': '<p>Content of first title</p> <p>Content of first title</p>', 'Second title': '<p>Content of second title</p>'}
解决方案
核心思路是先统一两种场景的目标节点:不管<strong>是独立存在还是嵌套在<p>内,都将其所在的同级节点纳入处理列表,再遍历节点收集后续内容。
适配代码如下:
from bs4 import BeautifulSoup def extract_title_content(html): soup = BeautifulSoup(html, 'html.parser') # 整理所有包含<strong>的目标节点 target_nodes = [] for strong_tag in soup.find_all('strong'): parent = strong_tag.parent # 若<strong>嵌套在<p>内,取<p>作为目标节点;否则直接取<strong> if parent.name == 'p' and parent.find('strong') == strong_tag: target_nodes.append(parent) else: target_nodes.append(strong_tag) result = {} for idx, current_node in enumerate(target_nodes): # 提取标题文本 title = current_node.get_text(strip=True) content_elements = [] next_sibling = current_node.next_sibling # 收集当前节点到下一个目标节点之间的所有有效内容 while next_sibling is not None: # 遇到下一个目标节点则停止 if next_sibling in target_nodes: break # 跳过空文本节点 if isinstance(next_sibling, str) and next_sibling.strip() == '': next_sibling = next_sibling.next_sibling continue # 将元素转为字符串存入列表 content_elements.append(str(next_sibling)) next_sibling = next_sibling.next_sibling # 拼接内容字符串 result[title] = ' '.join(content_elements) return result # 测试示例1 example1 = """<strong>First Title</strong><p>Content of first title</p><p>Content of first title</p><strong>Second title</strong><p>Content of second title</p>""" print(extract_title_content(example1)) # 测试示例2 example2 = """<p><strong>First Title</strong></p><p>Content of first title</p><p>Content of first title</p><p><strong>Second title</strong></p><p>Content of second title</p>""" print(extract_title_content(example2))
代码说明
- 先遍历所有
<strong>标签,判断其是否嵌套在<p>内,统一两种场景的处理节点,避免因层级差异导致失效。 - 遍历每个目标节点时,逐个检查后续兄弟节点,直到遇到下一个目标节点为止,期间跳过空文本节点,只收集有效元素。
- 将收集到的元素转为字符串后拼接,作为对应标题的内容值。
内容的提问来源于stack exchange,提问作者Rakesh Shetty
相关产品推荐
相关产品推荐

