You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:如何用BeautifulSoup解析h2及下属h3并提取指定h2文本

解决思路与代码示例

嘿,刚接触lxml和BeautifulSoup完全不用客气~新手入门都会有这样的小困惑,我来一步步帮你搞定!

一、获取指定h2标签的文本

你已经通过soup.find_all("h2")拿到了所有h2标签,接下来只需要遍历这些标签,判断里面<strong>的文本是否是你要找的目标即可:

from bs4 import BeautifulSoup

# 假设你已经把CSV中的HTML内容解析成了soup对象
for h2_tag in soup.find_all("h2"):
    # 提取h2下strong标签的文本,strip()去除多余空格和换行
    heading = h2_tag.find("strong").get_text(strip=True)
    # 判断是否是目标标题
    if heading in ["Powers and Abilities", "Weapons and Equipment"]:
        print(f"找到目标标题:{heading}")
        # 如果需要获取该标题下的具体内容,还可以用h2_tag.next_sibling等方法遍历后续节点

二、解析h2标签及其下属的h3标签

通常这类文档结构中,h3是h2的同级子章节(紧跟在h2之后,直到下一个h2出现),我们可以通过遍历h2的兄弟节点来收集对应的h3:

h2_list = soup.find_all("h2")
for h2 in h2_list:
    h2_title = h2.find("strong").get_text(strip=True)
    print(f"\n=== {h2_title} ===")
    
    # 从h2的下一个节点开始遍历
    current_node = h2.next_sibling
    while current_node:
        # 跳过空白文本节点(HTML中的换行、空格等)
        if isinstance(current_node, str) and current_node.strip() == "":
            current_node = current_node.next_sibling
            continue
        
        # 遇到下一个h2就停止当前h2的遍历
        if current_node.name == "h2":
            break
        
        # 如果是h3标签,提取文本
        if current_node.name == "h3":
            h3_title = current_node.get_text(strip=True)
            print(f"  - {h3_title}")
        
        # 移动到下一个节点
        current_node = current_node.next_sibling

小提醒

因为你是从CSV文件中读取数据,记得先通过csv模块把对应单元格里的HTML字符串提取出来,再用BeautifulSoup(html_str, "lxml")创建解析对象哦~

内容的提问来源于stack exchange,提问作者Lunalight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:19:39