You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将指定结构化数据转换为目标嵌套字典?解决重复值问题

问题:结构化数据转换为嵌套字典格式

原始数据

this_data = [{
    "Name": "Bluefox",
    "Sub Name": "Moonglow",
    "Time Series": "{'2022-07-06': 9.5, '2022-07-07': 7.2, '2022-07-08': 10.3}",
    "Probability": "{'2022-07-06': 0.2, '2022-07-07': 0.3, '2022-07-08': 0.5}",
    "Max Value": 466888785.24275005,
},{
    "Name": "Blackbird",
    "Sub Name": "Skylight",
    "Time Series": "{'2022-07-06': -16240599.020647092, '2022-07-07': -17984033.390385196}",
    "Probability": "{'2022-07-06': 0.6, '2022-07-07': 0.7}",
    "Max Value": 81509865.34667145,
},{
    "Name": "Bluefox",
    "Sub Name": "Skylight",
    "Time Series": "{'2022-07-06': -123000, '2022-07-07': -245100}",
    "Probability": "{'2022-07-06': 0.0, '2022-07-07': 0.0}",
    "Max Value": 90409417.34667145,
}]

目标格式

{
    'Bluefox': {
        'Moonglow': {
            'date': {
                '2022-07-06': {
                    'Time Series': 9.5,
                    'Probability': 0.2
                },
                '2022-07-07': {
                    'Time Series': 7.2,
                    'Probability': 0.3
                },
                '2022-07-08': {
                    'Time Series': 10.3,
                    'Probability': 0.5
                }
            },
            'Max Value': 466888785.24275005
        },
        'Skylight': {
            'date': {
                '2022-07-06': {
                    'Time Series': -123000,
                    'Probability': 0.0
                },
                '2022-07-07': {
                    'Time Series': -245100,
                    'Probability': 0.0
                }
            },
            'Max Value': 90409417.34667145
        }
    },
    'Blackbird': {
        'Skylight': {
            'date': {
                '2022-07-06': {
                    'Time Series': -16240599.020647092,
                    'Probability': 0.6
                },
                '2022-07-07': {
                    'Time Series': -17984033.390385196,
                    'Probability': 0.7
                }
            },
            'Max Value': 81509865.34667145
        }
    }
}

尝试的代码(存在值重复问题)

import json
import ast
from collections import defaultdict

entity_to_cp = {
    'Bluefox' : ['Moonglow', 'Skylight'],
    'Blackbird' : ['Skylight']
}


inner = defaultdict(list)
between = defaultdict(dict)
between2 = defaultdict(dict)
outer = defaultdict(dict)
for each_dict in this_data:
    for label, all_values in each_dict.items():
        if label == "Name":
            outer[all_values] = between
            cur_e = all_values
        if label == "Sub Name":
            between[all_values] = between2
            cur_cp = all_values
        
        try:
            if cur_cp in entity_to_cp[cur_e]:
                try:
                    all_values = ast.literal_eval(all_values)
                    for k,v in all_values.items():
                        print(k)
                        between2[k] = inner
                        inner[label].append(v)
                    # inner[label].append(all_values)
                except AttributeError as e:
                    print(e)
                except SyntaxError as e:
                    print(e)
                except ValueError as e:
                    print(e)
        except NameError as e:
            print(e)

问题原因

原代码使用了全局的defaultdict实例(inner、between、between2),所有数据条目都会复用这些字典,导致不同Name和Sub Name下的数据相互覆盖、重复。

解决方案代码

import ast

# 初始化最终结果字典
result = {}

for item in this_data:
    name = item["Name"]
    sub_name = item["Sub Name"]
    max_value = item["Max Value"]
    
    # 解析时间序列和概率的字符串为字典
    time_series = ast.literal_eval(item["Time Series"])
    probability = ast.literal_eval(item["Probability"])
    
    # 构建Name层级
    if name not in result:
        result[name] = {}
    
    # 构建Sub Name层级,初始化date和Max Value
    if sub_name not in result[name]:
        result[name][sub_name] = {
            "date": {},
            "Max Value": max_value
        }
    
    # 填充date下的日期数据
    date_dict = result[name][sub_name]["date"]
    for date in time_series:
        if date not in date_dict:
            date_dict[date] = {}
        date_dict[date]["Time Series"] = time_series[date]
        date_dict[date]["Probability"] = probability[date]

# 打印结果验证
import pprint
pprint.pprint(result)

代码解释

  1. 初始化结果字典:用普通字典存储最终嵌套结构,避免全局共享实例的问题。
  2. 遍历原始数据:逐个处理每个数据条目,提取Name、Sub Name、Max Value。
  3. 解析字符串字典:用ast.literal_eval把Time Series和Probability的字符串转为可操作的字典。
  4. 层级构建:
    • 检查Name是否在结果中,不存在则创建空字典。
    • 检查Sub Name是否在对应Name下,不存在则初始化包含date空字典和Max Value的结构。
  5. 填充日期数据:遍历每个日期,将对应的Time Series和Probability值存入date下的对应日期字典中。

内容的提问来源于stack exchange,提问作者DUDANF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 02:15:39