Python如何合并嵌套JSON?求替代R中tbl_json的方案
Python实现嵌套JSON扁平化(替代R的tbl_json)
你有如下嵌套JSON结构:
{ "index": "exp-000005", "type": "_doc", "score": 9.502488, "source": { "verb": "REPLIED", "timestamp": "2022-01-20T08:14:00+00:00", "in_context": { "screen_width": "3440", "screen_height": "1440", "build_version": "7235", "question": "Hallo", "request_time": "403", "status": "success" } } }
想要将其转为单层JSON:
{ "index": "exp-000005", "type": "_doc", "score": 9.502488, "verb": "REPLIED", "timestamp": "2022-01-20T08:14:00+00:00", "screen_width": "3440", "screen_height": "1440", "build_version": "7235", "question": "Hallo", "request_time": "403", "status": "success" }
之前在R中用tbl_json实现,以下是Python里的几种替代方案:
方法1:手动递归实现(无需第三方库)
自己写递归函数遍历嵌套字典,直接将所有键值对展平到单层:
def flatten_json(nested_dict, parent_key='', sep=''): items = [] for k, v in nested_dict.items(): new_key = f"{parent_key}{sep}{k}" if parent_key else k if isinstance(v, dict): items.extend(flatten_json(v, new_key, sep=sep).items()) else: items.append((new_key, v)) return dict(items) # 测试使用 nested_data = { "index": "exp-000005", "type": "_doc", "score": 9.502488, "source": { "verb": "REPLIED", "timestamp": "2022-01-20T08:14:00+00:00", "in_context": { "screen_width": "3440", "screen_height": "1440", "build_version": "7235", "question": "Hallo", "request_time": "403", "status": "success" } } } flat_data = flatten_json(nested_data) print(flat_data)
这个函数会自动处理所有嵌套层级,输出结果和目标格式完全匹配。
方法2:使用flatten_json第三方库
专门用于JSON扁平化的工具库,用法简洁:
- 先安装库:
pip install flatten_json
- 代码实现:
from flatten_json import flatten nested_data = { "index": "exp-000005", "type": "_doc", "score": 9.502488, "source": { "verb": "REPLIED", "timestamp": "2022-01-20T08:14:00+00:00", "in_context": { "screen_width": "3440", "screen_height": "1440", "build_version": "7235", "question": "Hallo", "request_time": "403", "status": "success" } } } # 设置sep=''避免键名添加分隔符,匹配目标格式 flat_data = flatten(nested_data, sep='') print(flat_data)
方法3:使用pandas的json_normalize
适合处理批量JSON数据,还能直接转换为DataFrame格式:
import pandas as pd nested_data = { "index": "exp-000005", "type": "_doc", "score": 9.502488, "source": { "verb": "REPLIED", "timestamp": "2022-01-20T08:14:00+00:00", "in_context": { "screen_width": "3440", "screen_height": "1440", "build_version": "7235", "question": "Hallo", "request_time": "403", "status": "success" } } } # 展平数据并转为字典 df = pd.json_normalize(nested_data, sep='') flat_data = df.to_dict(orient='records')[0] print(flat_data)
内容的提问来源于stack exchange,提问作者threxx
相关产品推荐
相关产品推荐

