You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python JSON转CSV遇嵌套层级问题,循环数据映射出错求助

解决JSON-stat格式数据转CSV的索引映射问题

你的核心问题是错误地将计算出的数字索引直接作为label字典的键。JSON-stat格式中,dimension下的category.label的键是维度成员的编码(比如STATISTIC的label键是"HSA06A01"这类字符串),而不是连续的数字;category.index才是成员在值数组中的顺序映射,你需要先通过index获取正确的成员编码,再用编码去取对应的label。

以下是修正后的完整代码:

import json
import csv
 
local_file_path = "/Users/piuno/OneDrive - Asociación Club Taller Cultural/full_data_housing.json"

# 从文件加载JSON数据
with open(local_file_path, 'r') as json_file:
    data = json.load(json_file)
 
# 写入CSV文件
with open('output.csv', mode='w', newline='') as file:
    writer = csv.writer(file)
    # 写入表头
    header = ['Statistic', 'Quarter', 'Area', 'Value']
    writer.writerow(header)
    
    # 获取维度的成员顺序(与值数组的排列顺序完全对应)
    statistic_index_order = list(data['dataset']['dimension']['STATISTIC']['category']['index'].values())
    statistic_label = data['dataset']['dimension']['STATISTIC']['category']['label']
    
    quarter_index_order = list(data['dataset']['dimension']['TLIST(Q1)']['category']['index'].values())
    quarter_label = data['dataset']['dimension']['TLIST(Q1)']['category']['label']
    
    area_index_order = list(data['dataset']['dimension']['C02343V02817']['category']['index'].values())
    area_label = data['dataset']['dimension']['C02343V02817']['category']['label']
    
    values = data['dataset']['value']
    
    # 计算每个维度的成员数量,用于索引步长计算
    quarter_count = len(quarter_index_order)
    area_count = len(area_index_order)
    
    for i in range(len(values)):
        if values[i] is not None:
            # 计算当前值对应的三个维度的位置索引
            statistic_pos = i // (quarter_count * area_count)
            quarter_pos = (i % (quarter_count * area_count)) // area_count
            area_pos = i % area_count
            
            # 通过位置获取成员编码,再匹配对应的标签名称
            statistic_name = statistic_label[statistic_index_order[statistic_pos]]
            quarter_name = quarter_label[quarter_index_order[quarter_pos]]
            area_name = area_label[area_index_order[area_pos]]
            
            row = [statistic_name, quarter_name, area_name, values[i]]
            writer.writerow(row)

关键修正说明

  • 先通过category.index.values()提取维度成员的顺序列表,这个列表的顺序和value数组的排列逻辑完全一致
  • 用维度成员数量计算索引步长,避免重复读取嵌套结构,提升代码可读性
  • 通过位置从顺序列表中获取成员编码,再用编码去label字典中匹配名称,彻底解决索引不匹配的问题

内容的提问来源于stack exchange,提问作者Andrew Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:44:52