如何优化Python解析AWS API未知结构JSON的输出格式?
解决方案:扁平化AWS API JSON并移除层级前缀与列表索引
核心思路
要实现需求,核心是遍历嵌套JSON时只保留最底层的键名,自动忽略中间层级节点和列表索引。需重点处理嵌套对象、嵌套数组两种场景,同时要考虑键名冲突的处理逻辑(比如同名键值的合并或覆盖)。
自定义实现(Python递归版)
以下是适配任意AWS API JSON结构的递归实现,逻辑简洁且兼容性强:
def flatten_json_simple(data): result = {} def traverse(current, current_key): if isinstance(current, dict): for k, v in current.items(): # 直接传递当前键,跳过上层路径 traverse(v, k) elif isinstance(current, list): for item in current: # 遍历列表元素,忽略索引直接处理 traverse(item, current_key) else: # 处理键名冲突:若键已存在则转为列表存储 if current_key in result: if not isinstance(result[current_key], list): result[current_key] = [result[current_key]] result[current_key].append(current) else: result[current_key] = current traverse(data, "") return result # 测试EC2接口模拟数据 sample_ec2_response = { "NetworkInterfaces": [ { "Groups": [ {"GroupName": "web-sg", "GroupId": "sg-123"}, {"GroupName": "db-sg", "GroupId": "sg-456"} ], "PrivateIpAddress": "10.0.0.7" } ] } print(flatten_json_simple(sample_ec2_response)) # 输出:{'GroupName': ['web-sg', 'db-sg'], 'GroupId': ['sg-123', 'sg-456'], 'PrivateIpAddress': '10.0.0.7'}
关键细节说明
- 递归遍历逻辑:通过递归穿透所有嵌套层级,只保留最终的键值对,自动跳过中间路径和列表索引。
- 冲突处理:示例中遇到同名键会将值合并为列表,你可根据需求调整(比如覆盖、添加后缀区分等)。
- 兼容性:支持AWS API返回的任意复杂嵌套结构,包括多层对象与数组的组合。
更优实现方式探讨
1. 迭代式遍历(避免栈溢出)
如果处理超大规模JSON,递归可能触发栈溢出,可改用栈实现迭代式遍历,逻辑与递归一致但稳定性更强:
def flatten_json_iterative(data): result = {} stack = [(data, "")] while stack: current, key = stack.pop() if isinstance(current, dict): for k, v in current.items(): stack.append((v, k)) elif isinstance(current, list): for item in current: stack.append((item, key)) else: if key in result: if not isinstance(result[key], list): result[key] = [result[key]] result[key].append(current) else: result[key] = current return result
2. 第三方库简化实现(可选)
若允许使用第三方库,pandas的json_normalize可快速扁平化JSON,再通过列名处理移除层级前缀:
import pandas as pd def flatten_with_pandas(data): # 先扁平化JSON df = pd.json_normalize(data, sep='.') # 提取列名的最后一段作为新键 column_mapping = {col: col.split('.')[-1] for col in df.columns} df = df.rename(columns=column_mapping) # 合并同名列的重复值 for col in df.columns: if df[col].count() > 1: df[col] = df[col].dropna().tolist() return df.to_dict('records')[0]
注意:这种方式在处理复杂数组嵌套时,需要额外调整json_normalize的参数,灵活性略逊于自定义实现。
3. 可配置化扩展
如果需要在部分场景保留特定层级,可给函数添加参数(比如keep_prefixes列表),让逻辑更灵活,满足不同业务需求。
内容的提问来源于stack exchange,提问作者Effie
相关产品推荐
相关产品推荐

