You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup解析维基百科表格生成递归嵌套JSON结构的实现困境与求助

基于BeautifulSoup解析维基百科表格生成递归嵌套JSON结构的实现困境与求助

我目前正在尝试把维基百科里的一个表格转换成嵌套的JSON对象,核心需求是将每个条目归类到对应的父分类下。举个例子,理想的输出结构应该是这样的:

{
  "name": "Mining, quarrying, and oil and gas extraction",
  "value": 593300,
  "children": [
    {
      "name": "Oil and gas extraction",
      "value": 118500,
      "children": []
    },
    {
      "name": "Mining, except oil and gas",
      "value": 189700,
      "children": [
        {
          "name": "Coal mining",
          "value": 44100,
          "children": []
        },
        {
          "name": "Metal ore mining",
          "value": 44100,
          "children": []
        },
        {
          "name": "Nonmetallic mineral mining and quarrying",
          "value": 102700,
          "children": []
        }
      ]
    },
    {
      "name": "Support activities for mining",
      "value": 290100,
      "children": []
    }
  ]
}

我用BeautifulSoup来解析页面,能把表格行遍历成列表,并且通过<td>标签的style: padding属性来判断缩进级别*(为了让索引方法从0开始,我手动给本地网页文件里第一个非表头的<td>添加了style="padding-left: 0em")。但现在的问题是,我没法把缩进级别转换成有意义的父子层级关系。

我尝试用递归的方式来处理这个列表,写了下面这段代码:

import bs4

f = open("./webpage.html", "r", encoding="utf8")
soup = bs4.BeautifulSoup(f, "html.parser")
table = soup.find( "table", attrs={"class": "collapsible wikitable mw-collapsible mw-made-collapsible"} )
rows = table.find_all("tr")

stash = []

def recurse(start_idx):
    for idx, row in enumerate(rows[start_idx:]):
        cols = row.find_all("td")
        # skip headers
        if len(cols) == 0:
            continue
        if not cols[0].attrs['style']:
            continue
        indent = int(cols[0].attrs["style"].strip(';')[-3])
        entry = {
            "name": cols[0].text,
            "value": cols[1].text,
            "indent": indent,
            "children": [],
        }
        while stash:
            parent = stash[-1]
            for inner_idx, inner_row in enumerate(rows[start_idx + idx + 1:]):
                inner_cols = inner_row.find_all("td")
                inner_indent = int(inner_cols[0].attrs["style"].strip(';')[-3])
                inner_entry = {
                    "name": inner_cols[0].text,
                    "value": inner_cols[1].text,
                    "indent": inner_indent,
                    "children": [],
                }
                if indent == parent["indent"]:
                    child = stash.pop()
                    stash[-1]["children"].append(child)
                    stash.append(entry)
                if parent["indent"] - indent == -1:
                    stash.append(entry)
                    # add row to stash
                    recurse(start_idx + idx + 1)
                    # calculate children of row
                    parent["children"].append(entry)
                # add row with children to row parent
                if indent < parent["indent"]:
                    return
        stash.append(entry)

recurse(0)

现在这个方法能把直接子节点正确添加到父节点里,但没法正确处理子节点的子节点(也就是更深层级的嵌套)。我感觉这个问题应该有个更巧妙的数据结构方案,但我查了不少资料都没找到头绪,希望能得到大家的指导和建议!

*注:我需要手动给.html文件里第一个非表头的<td>添加style="padding-left: 0em",这样我的索引方法才能从0开始。

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 08:34:36