如何提取字典中的嵌套字典并转换为单层字典?
如何将包含嵌套字典的Python字典扁平化为单层键值对?
我有一个字典,部分键值对的值是基础类型,部分是嵌套字典。数据示例如下:
original_dict = { 'amount': 123, 'baseUnit': 'test', 'currency': {'code': 'EUR'}, 'dimensions': { 'height': {'iri': 'http://www.example.com/data/measurement-height-12345', 'unitOfMeasure': 'm', 'value': 23}, 'length': {'iri': 'http://www.example.com/data/measurement-length-12345', 'unitOfMeasure': 'm', 'value': 8322}, 'volume': {'unitOfMeasure': '', 'value': 0}, 'weight': {'iri': 'http://www.example.com/data/measurement-weight-12345', 'unitOfMeasure': 'KG', 'value': 23}, 'width': {'iri': 'http://www.example.com/data/measurement-width-12345', 'unitOfMeasure': 'm', 'value': 1} }, 'exportListNumber': '1234', 'iri': 'http://www.example.com/data/material-12345', 'number': '12345', 'orderUnit': 'sdf', 'producerFormattedPID': '12345', 'producerID': 'example', 'producerNonFormattedPID': '12345', 'stateID': 'm70', 'typeID': 'FERT' }
我需要将嵌套的部分(比如dimensions和price,示例中price未体现但需求类似)展开,得到一个仅包含扁平键值对的字典。例如,类似price: {'currency': {'code': 'EUR'}, 'amount': 123}的结构要转换为{'pricecurrencycode':'EUR','priceamount':123};dimensions下的所有嵌套内容也要全部提取,方便后续转换为DataFrame。
方法一:通用递归扁平化函数
如果需要对所有嵌套字典进行通用扁平化,用下划线拼接父键和子键,可以用递归函数实现:
def flatten_dict(d, parent_key='', sep='_'): items = [] for k, v in d.items(): new_key = f"{parent_key}{sep}{k}" if parent_key else k if isinstance(v, dict): items.extend(flatten_dict(v, new_key, sep=sep).items()) else: items.append((new_key, v)) return dict(items) # 调用函数 flattened = flatten_dict(original_dict) print(flattened)
输出结果示例(部分):
{ 'amount': 123, 'baseUnit': 'test', 'currency_code': 'EUR', 'dimensions_height_iri': 'http://www.example.com/data/measurement-height-12345', 'dimensions_height_unitOfMeasure': 'm', 'dimensions_height_value': 23, # ... 其他展开后的键值对 }
方法二:定制化符合特定命名规则的扁平化
如果需要像你要求的那样,把price相关的字段命名为pricecurrencycode、priceamount,或者对dimensions有特殊命名规则,可以针对性处理:
def custom_flatten(d): flattened = {} # 先处理非嵌套的基础字段 for k, v in d.items(): if k not in ['currency', 'dimensions', 'price']: flattened[k] = v # 处理currency(对应price相关的命名) if 'currency' in d: flattened['pricecurrencycode'] = d['currency']['code'] if 'amount' in d: flattened['priceamount'] = d['amount'] # 处理dimensions if 'dimensions' in d: for dim_name, dim_data in d['dimensions'].items(): for sub_k, sub_v in dim_data.items(): # 可自定义命名规则,比如dim_name + 下划线 + sub_k new_key = f"{dim_name}_{sub_k}" flattened[new_key] = sub_v return flattened # 调用函数 custom_flattened = custom_flatten(original_dict) print(custom_flattened)
输出结果示例(部分):
{ 'baseUnit': 'test', 'exportListNumber': '1234', 'iri': 'http://www.example.com/data/material-12345', # ... 其他基础字段 'pricecurrencycode': 'EUR', 'priceamount': 123, 'height_iri': 'http://www.example.com/data/measurement-height-12345', 'height_unitOfMeasure': 'm', 'height_value': 23, # ... 其他dimensions展开的字段 }
方法三:使用pandas的json_normalize(直接转DataFrame)
如果最终目的是转换为DataFrame,可以直接用pandas的json_normalize函数,它会自动处理嵌套字典:
import pandas as pd df = pd.json_normalize(original_dict, sep='_') print(df.columns) # 输出的列名就是扁平化后的键,比如'currency_code'、'dimensions_height_iri'等
这种方法无需手动编写扁平化函数,直接生成符合需求的DataFrame。
内容的提问来源于stack exchange,提问作者kitten_world
相关产品推荐
相关产品推荐

