基于BeautifulSoup解析维基百科表格生成递归嵌套JSON结构的实现困境与求助
基于BeautifulSoup解析维基百科表格生成递归嵌套JSON结构的实现困境与求助
我目前正在尝试把维基百科里的一个表格转换成嵌套的JSON对象,核心需求是将每个条目归类到对应的父分类下。举个例子,理想的输出结构应该是这样的:
{ "name": "Mining, quarrying, and oil and gas extraction", "value": 593300, "children": [ { "name": "Oil and gas extraction", "value": 118500, "children": [] }, { "name": "Mining, except oil and gas", "value": 189700, "children": [ { "name": "Coal mining", "value": 44100, "children": [] }, { "name": "Metal ore mining", "value": 44100, "children": [] }, { "name": "Nonmetallic mineral mining and quarrying", "value": 102700, "children": [] } ] }, { "name": "Support activities for mining", "value": 290100, "children": [] } ] }
我用BeautifulSoup来解析页面,能把表格行遍历成列表,并且通过<td>标签的style: padding属性来判断缩进级别*(为了让索引方法从0开始,我手动给本地网页文件里第一个非表头的<td>添加了style="padding-left: 0em")。但现在的问题是,我没法把缩进级别转换成有意义的父子层级关系。
我尝试用递归的方式来处理这个列表,写了下面这段代码:
import bs4 f = open("./webpage.html", "r", encoding="utf8") soup = bs4.BeautifulSoup(f, "html.parser") table = soup.find( "table", attrs={"class": "collapsible wikitable mw-collapsible mw-made-collapsible"} ) rows = table.find_all("tr") stash = [] def recurse(start_idx): for idx, row in enumerate(rows[start_idx:]): cols = row.find_all("td") # skip headers if len(cols) == 0: continue if not cols[0].attrs['style']: continue indent = int(cols[0].attrs["style"].strip(';')[-3]) entry = { "name": cols[0].text, "value": cols[1].text, "indent": indent, "children": [], } while stash: parent = stash[-1] for inner_idx, inner_row in enumerate(rows[start_idx + idx + 1:]): inner_cols = inner_row.find_all("td") inner_indent = int(inner_cols[0].attrs["style"].strip(';')[-3]) inner_entry = { "name": inner_cols[0].text, "value": inner_cols[1].text, "indent": inner_indent, "children": [], } if indent == parent["indent"]: child = stash.pop() stash[-1]["children"].append(child) stash.append(entry) if parent["indent"] - indent == -1: stash.append(entry) # add row to stash recurse(start_idx + idx + 1) # calculate children of row parent["children"].append(entry) # add row with children to row parent if indent < parent["indent"]: return stash.append(entry) recurse(0)
现在这个方法能把直接子节点正确添加到父节点里,但没法正确处理子节点的子节点(也就是更深层级的嵌套)。我感觉这个问题应该有个更巧妙的数据结构方案,但我查了不少资料都没找到头绪,希望能得到大家的指导和建议!
*注:我需要手动给.html文件里第一个非表头的<td>添加style="padding-left: 0em",这样我的索引方法才能从0开始。
内容来源于stack exchange
相关产品推荐
相关产品推荐

