如何将指定结构化数据转换为目标嵌套字典?解决重复值问题
问题:结构化数据转换为嵌套字典格式
原始数据
this_data = [{ "Name": "Bluefox", "Sub Name": "Moonglow", "Time Series": "{'2022-07-06': 9.5, '2022-07-07': 7.2, '2022-07-08': 10.3}", "Probability": "{'2022-07-06': 0.2, '2022-07-07': 0.3, '2022-07-08': 0.5}", "Max Value": 466888785.24275005, },{ "Name": "Blackbird", "Sub Name": "Skylight", "Time Series": "{'2022-07-06': -16240599.020647092, '2022-07-07': -17984033.390385196}", "Probability": "{'2022-07-06': 0.6, '2022-07-07': 0.7}", "Max Value": 81509865.34667145, },{ "Name": "Bluefox", "Sub Name": "Skylight", "Time Series": "{'2022-07-06': -123000, '2022-07-07': -245100}", "Probability": "{'2022-07-06': 0.0, '2022-07-07': 0.0}", "Max Value": 90409417.34667145, }]
目标格式
{ 'Bluefox': { 'Moonglow': { 'date': { '2022-07-06': { 'Time Series': 9.5, 'Probability': 0.2 }, '2022-07-07': { 'Time Series': 7.2, 'Probability': 0.3 }, '2022-07-08': { 'Time Series': 10.3, 'Probability': 0.5 } }, 'Max Value': 466888785.24275005 }, 'Skylight': { 'date': { '2022-07-06': { 'Time Series': -123000, 'Probability': 0.0 }, '2022-07-07': { 'Time Series': -245100, 'Probability': 0.0 } }, 'Max Value': 90409417.34667145 } }, 'Blackbird': { 'Skylight': { 'date': { '2022-07-06': { 'Time Series': -16240599.020647092, 'Probability': 0.6 }, '2022-07-07': { 'Time Series': -17984033.390385196, 'Probability': 0.7 } }, 'Max Value': 81509865.34667145 } } }
尝试的代码(存在值重复问题)
import json import ast from collections import defaultdict entity_to_cp = { 'Bluefox' : ['Moonglow', 'Skylight'], 'Blackbird' : ['Skylight'] } inner = defaultdict(list) between = defaultdict(dict) between2 = defaultdict(dict) outer = defaultdict(dict) for each_dict in this_data: for label, all_values in each_dict.items(): if label == "Name": outer[all_values] = between cur_e = all_values if label == "Sub Name": between[all_values] = between2 cur_cp = all_values try: if cur_cp in entity_to_cp[cur_e]: try: all_values = ast.literal_eval(all_values) for k,v in all_values.items(): print(k) between2[k] = inner inner[label].append(v) # inner[label].append(all_values) except AttributeError as e: print(e) except SyntaxError as e: print(e) except ValueError as e: print(e) except NameError as e: print(e)
问题原因
原代码使用了全局的defaultdict实例(inner、between、between2),所有数据条目都会复用这些字典,导致不同Name和Sub Name下的数据相互覆盖、重复。
解决方案代码
import ast # 初始化最终结果字典 result = {} for item in this_data: name = item["Name"] sub_name = item["Sub Name"] max_value = item["Max Value"] # 解析时间序列和概率的字符串为字典 time_series = ast.literal_eval(item["Time Series"]) probability = ast.literal_eval(item["Probability"]) # 构建Name层级 if name not in result: result[name] = {} # 构建Sub Name层级,初始化date和Max Value if sub_name not in result[name]: result[name][sub_name] = { "date": {}, "Max Value": max_value } # 填充date下的日期数据 date_dict = result[name][sub_name]["date"] for date in time_series: if date not in date_dict: date_dict[date] = {} date_dict[date]["Time Series"] = time_series[date] date_dict[date]["Probability"] = probability[date] # 打印结果验证 import pprint pprint.pprint(result)
代码解释
- 初始化结果字典:用普通字典存储最终嵌套结构,避免全局共享实例的问题。
- 遍历原始数据:逐个处理每个数据条目,提取
Name、Sub Name、Max Value。 - 解析字符串字典:用
ast.literal_eval把Time Series和Probability的字符串转为可操作的字典。 - 层级构建:
- 检查
Name是否在结果中,不存在则创建空字典。 - 检查
Sub Name是否在对应Name下,不存在则初始化包含date空字典和Max Value的结构。
- 检查
- 填充日期数据:遍历每个日期,将对应的
Time Series和Probability值存入date下的对应日期字典中。
内容的提问来源于stack exchange,提问作者DUDANF
相关产品推荐
相关产品推荐

