如何自动化归一化分形结构的嵌套JSON并生成完整语句
解决方案:递归深度优先遍历生成完整语句
你的需求本质上是遍历树形结构的所有叶子节点路径,把路径上的label依次拼接成句子。这种场景用**递归深度优先搜索(DFS)**最适合,完全不需要预定义结构,能自动适配各种层级和宽度的JSON文件。
核心实现代码
def extract_sentences(node, current_path=None): # 初始化当前路径,第一次调用时为空列表 if current_path is None: current_path = [] # 把当前节点的label加入路径 current_path = current_path + [node['label']] # 如果当前节点没有children,说明走到了路径尽头,返回拼接好的句子 if 'children' not in node or not node['children']: return [' '.join(current_path)] # 否则遍历所有子节点,递归收集所有子路径的句子 sentences = [] for child in node['children']: sentences.extend(extract_sentences(child, current_path)) return sentences # 处理你的示例数据 test = [ { "label": "I", "children": [ { "label": "want", "children": [ { "label": "a", "children": [ {"label": "coffee"}, {"label": "big", "children": [{"label": "piece of cake"}]}, ], } ], }, {"label": "need", "children": [{"label": "time"}]}, {"label": "like", "children": [{"label": "italian", "children": [{"label": "pizza"}]}]}, ], }, { "label": "We", "children": [ {"label": "are", "children": [{"label": "ok"}]}, {"label": "will", "children": [{"label": "rock you"}]}, ], }, ] # 遍历根节点列表,收集所有句子 all_sentences = [] for root in test: all_sentences.extend(extract_sentences(root)) print(all_sentences)
输出结果
[ 'I want a coffee', 'I want a big piece of cake', 'I need time', 'I like italian pizza', 'We are ok', 'We will rock you' ]
为什么这个方法能解决你的问题?
- 完全动态适配:不管你的JSON层级有多深、每个节点的子节点数量差异多大,递归都会自动遍历所有可能的路径,不需要提前定义任何结构参数(比如
json_normalize的meta/record_path)。 - 路径跟踪清晰:每次递归都会带着当前已经拼接的路径,遇到叶子节点(没有children的节点)就直接生成完整句子,完美对应你要的"所有路径"逻辑,和
os.walk遍历文件路径的思路一致。 - 代码简洁易维护:核心逻辑只有十几行,后续如果需要调整拼接规则(比如用其他分隔符),只需要修改
' '.join(current_path)这一行即可。
对比你之前尝试的方案
- 比
pandas.json_normalize更灵活:不需要提前分析结构,适用于100多个不同结构的JSON文件。 - 比
jsonpath_ng更精准:不仅能提取所有label,还能保留层级关系,正确拼接成完整语句。 - 比扁平化/层级元组方法更直接:不需要处理复杂的索引拆分,递归天然就能处理树形结构的路径传递。
如果你的JSON文件里存在children为空数组的情况(比如{"label": "test", "children": []}),这个函数也能正确处理,会把"test"作为单独的句子返回。
内容的提问来源于stack exchange,提问作者R1mk4
相关产品推荐
相关产品推荐

