如何使用Python展平嵌套层级JSON数据以适配BigQuery加载
嵌套部门JSON扁平化转换方案
问题背景
现有如下递归嵌套结构的部门组织JSON数据:
{ "department":"Data & Analytics", "child":[ { "department":"Data Enginnering", "child": [ {"department":"AWS Squad"}, {"department":"GCP Squad"}, "..so..", "..on..", "..so..", "..forth.." ] }, { "department":"Data Science" } ] }
为了将该数据导入BigQuery,需要将嵌套结构转换为扁平化列表结构:每个部门作为列表中的独立对象,child字段仅存储该部门直接下属的部门名称数组,无下属部门的对象不携带child字段,目标结构如下:
[ { "department":"Data & Analytics", "child":["Data Enginnering", "Data Science"] }, { "department":"Data Enginnering", "child":["AWS Squad", "GCP Squad"] }, { "department":"Data Science" }, { "department": "AWS Squad" }, { "department": "GCP Squad" } ]
实现代码
使用Python基于广度优先遍历实现任意深度嵌套结构的转换,代码如下:
import json from collections import deque def transform_dept_data(root_dept): result = [] process_queue = deque([root_dept]) while process_queue: current = process_queue.popleft() current_dept = {"department": current["department"]} child_depts = current.get("child", []) if child_depts: current_dept["child"] = [dept["department"] for dept in child_depts] process_queue.extend(child_depts) result.append(current_dept) return result # 调用示例 if __name__ == "__main__": # 替换为你的原始JSON文件读取逻辑即可 with open("your_raw_dept.json", "r", encoding="utf-8") as f: raw_data = json.load(f) output = transform_dept_data(raw_data) # 转换后的数据可以直接导出为JSON文件供BigQuery导入 with open("flattened_dept_for_bq.json", "w", encoding="utf-8") as f: json.dump(output, f, indent=4, ensure_ascii=False)
逻辑说明
- 遍历逻辑支持任意深度的部门嵌套,不存在层级限制
- 转换后的数据结构完全匹配BigQuery的表结构要求:
department为字符串类型,child为字符串数组类型(REPEATED模式) - 无下属部门的节点自动省略
child字段,不会产生空值字段干扰导入
内容的提问来源于stack exchange,提问作者luisvenezian
相关产品推荐
相关产品推荐

