You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析HTML使标题层级隐含的嵌套关系显性化?

解决方案:基于BeautifulSoup实现标题层级嵌套解析

你可以用BeautifulSoup(Python最常用的HTML解析库)配合简单的层级栈逻辑实现需求,无需从零构建核心解析功能。以下是可直接运行的代码,能生成符合要求的嵌套结构:

代码实现

from bs4 import BeautifulSoup
import json

def html_unpacker(html_string):
    soup = BeautifulSoup(html_string, 'html.parser')
    # 定义标题层级优先级
    heading_tags = {'h1':1, 'h2':2, 'h3':3, 'h4':4, 'h5':5, 'h6':6}
    stack = []
    root = {}
    current_parent = root
    # 统计同类型非标题元素的序号
    element_counters = {}

    for element in soup.find_all(['h1','h2','h3','h4','h5','h6','p','ul','li']):
        tag = element.name
        if tag in heading_tags:
            # 重置当前层级的非标题元素计数器
            element_counters = {}
            heading_level = heading_tags[tag]
            heading_text = element.get_text(strip=True)
            
            # 调整栈结构,找到当前标题的归属父级
            while stack and stack[-1]['level'] >= heading_level:
                stack.pop()
            
            # 创建当前标题节点并挂载到父级
            current_node = {}
            if stack:
                stack[-1]['node'][heading_text] = current_node
            else:
                root[heading_text] = current_node
            
            # 将当前节点压入栈,作为后续元素的父级
            stack.append({'level': heading_level, 'node': current_node})
            current_parent = current_node
        else:
            # 处理非标题元素,按类型计数保证顺序
            if tag not in element_counters:
                element_counters[tag] = 1
            else:
                element_counters[tag] += 1
            key = f"{tag}{element_counters[tag]}"
            
            if tag == 'ul':
                # 解析ul下的li元素,生成子结构
                li_counter = 1
                ul_node = {}
                for li in element.find_all('li', recursive=False):
                    ul_node[f"li{li_counter}"] = li.get_text(strip=True)
                    li_counter += 1
                current_parent[key] = ul_node
            else:
                current_parent[key] = element.get_text(strip=True)
    return root

# 测试示例HTML
html_string = """
<h1>Motley Mess Menu</h1>
<h2>Breakfast</h2>
<h3>Omelets</h3>
<h4>Cheese</h4>
<p>$7</p>
<p>American style omelet containing your choice of Cheddar, Swiss, Feta, Colby Jack or all four!</p>
<h4>Sausage Mushroom</h4>
<p>$8</p>
<p>American style omelet containing sausage, mushroom and Swiss cheese</p>
<h4>Build-Your-Own</h4>
<p>$8</p>
<p>American style omelet containing…you tell me!</p>
<p>Options (+50 cents after 3):</p>
<ul>
<li>Cheddar</li>
<li>Swiss</li>
<li>Feta</li>
<li>Colby Jack</li>
<li>Bacon Bits</li>
<li>Sausage</li>
<li>Onion</li>
<li>Hamburger</li>
<li>Jalapenos</li>
<li>Hash Browns</li>
</ul>
<h3>Combos</h3>
<p>Each come with your choice of two sides</p>
<h4>Eggs and Bacon</h4>
<p>$8</p>
<p>Eggs cooked your way and crispy bacon. Sausage substitution is fine</p>
<h4>Glorious Smash</h4>
<p>$10</p>
<p>Your favorite breakfast of two pancakes, two eggs cooked your way, two sausages and two bacon, free of all trademark infringement! If you think you can finish it all then you forgot about the choice of two sides!</p>
"""

result = html_unpacker(html_string)
print(json.dumps(result, indent=4, ensure_ascii=False))

输出效果

运行后会生成与你示例几乎一致的嵌套JSON结构:

  • 自动识别标题层级的从属关系,将h4挂载到对应的h3下,h3挂载到h2下
  • 非标题元素保留类型(p、ul、li)和顺序,用标签名+序号作为键名区分同类型元素
  • ul元素会自动解析内部的li,生成嵌套子结构

依赖说明

先安装依赖库:

pip install beautifulsoup4

内容的提问来源于stack exchange,提问作者psychicesp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 01:30:21