You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用BeautifulSoup:获取匹配指定内容的标签区间

解决BeautifulSoup抓取板块内容的问题

1. 定位匹配的<h1>元素

由于部分<h1>包含<img>标签,直接用.get_text()会混入图片相关文本,需排除<img>节点后提取纯文本进行匹配:

def find_target_h1(soup, target_name):
    for h1 in soup.find_all('h1'):
        # 提取<h1>中排除<img>的文本,去除首尾空白
        clean_text = ''.join([node.strip() for node in h1.contents if node.name != 'img']).strip()
        if clean_text == target_name:
            return h1
    return None  # 未找到匹配的<h1>

2. 获取目标<h1>到下一个<h1>(或文档末尾)的内容

利用BeautifulSoup的.next_siblings遍历目标<h1>的后续同级节点,遇到下一个<h1>就终止,可通过参数控制是否包含目标<h1>:

def get_section_content(soup, target_name, include_h1=True):
    target_h1 = find_target_h1(soup, target_name)
    if not target_h1:
        return None
    
    content_parts = []
    if include_h1:
        content_parts.append(target_h1)
    
    # 遍历后续兄弟节点,收集内容直到下一个<h1>
    for sibling in target_h1.next_siblings:
        if sibling.name == 'h1':
            break
        # 过滤无意义的空文本节点(比如换行、空格)
        if sibling.name or str(sibling).strip():
            content_parts.append(sibling)
    
    # 将节点转为HTML字符串返回
    return ''.join(str(part) for part in content_parts)

使用示例

# 假设soup是已解析的BeautifulSoup对象,name是目标板块名称
section_html = get_section_content(soup, name, include_h1=False)
if section_html:
    print(section_html)
else:
    print("未找到匹配的板块")

关键说明

  • .next_siblings会遍历当前节点后的所有同级节点,包括文本节点,需过滤空文本避免冗余内容。
  • 匹配文本时用.strip()处理首尾空白,防止网页中的空格、换行导致匹配失败。
  • 若需要保留BeautifulSoup节点对象,直接返回content_parts列表即可,无需转为字符串。

内容的提问来源于stack exchange,提问作者DonielF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 13:01:09