You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理JSON表格数据时列表出现重复字典值的原因排查与解决咨询

问题:Python处理JSON表格数据时列表出现重复字典值

我在处理一份包含版本发布统计数据的JSON格式表格数据时,编写了Python代码尝试将每行数据转换为对应字段的键值对字典,并将这些字典存储到列表中,但最终输出的列表中出现了重复的字典值。以下是相关信息:

JSON数据

{"name":"Sheet1!A1:M26","rows":[[{"v":"Date"},{"v":"Release"},{"v":"Functional Ench"},{"v":"PDBs"},{"v":"Total Changes Deployed"},{"v":"Post HCL Exemption"},{"v":"% of Post HCL compared to total changes"},{"v":"Control Room Items"},{"v":"Control Room items fixed (During hypercare week)"},{"v":"% of Control Room compared to total changes"},{"v":"% fixed vs Open items (During hypercare week)"},{"v":"Number of Hotfixes"},{"v":"Hotfix Prod\nDeploy Date"}],[{"v":44625,"fv":"3/5/2022"},{"v":"15.1.1"},{"v":16,"fv":"16"},{"v":26,"fv":"26"},{"v":42,"fv":"42"},{"v":2,"fv":"2"},{},{},{},{},{},{},{}],[{"v":44597,"fv":"2/5/2022"},{"v":"15.1.0"},{"v":41,"fv":"41"},{"v":5,"fv":"5"},{"v":46,"fv":"46"},{"v":3,"fv":"3"},{"v":6,"fv":"6"},{"v":3,"fv":"3"},{"v":3,"fv":"3"},{"v":13,"fv":"13"},{"v":50,"fv":"50"},{"v":3,"fv":"3"},{"v":44606,"fv":"14-Feb-22"}],...(其余行省略)]

原Python代码

import json
import pprint
fp = open('C:\\Users\\ADMIN\\Desktop\\test\\asd-Copy.txt')
data = json.load(fp)
a = {}
obj =list()
for j in range(1,len(data['rows'])):
    for x in range(0,len(data['rows'][j])):
        if len(data['rows'][j][x]) == 1:
            a[data['rows'][0][x]['v']]= data['rows'][j][x]['v']
        if len(data['rows'][j][x]) == 2:
            a[data['rows'][0][x]['v']]= data['rows'][j][x]['fv']
        if len(data['rows'][j][x]) == 0:
            a[data['rows'][0][x]['v']]= data['rows'][j][x]
    pprint.pprint(a)
    obj.append(a)
pprint.pprint(obj)

期望输出

[
    {'% fixed vs Open items (During hypercare week)': {}, '% of Control Room compared to total changes': {}, '% of Post HCL compared to total changes': {}, 'Control Room Items': {}, 'Control Room items fixed (During hypercare week)': {}, 'Date': '3/5/2022', 'Functional Ench': '16', 'Hotfix Prod\nDeploy Date': {}, 'Number of Hotfixes': {}, 'PDBs': '26', 'Post HCL Exemption': '2', 'Release': '15.1.1', 'Total Changes Deployed': '42'},
    {'% fixed vs Open items (During hypercare week)': '50', '% of Control Room compared to total changes': '13', '% of Post HCL compared to total changes': '6', 'Control Room Items': '3', 'Control Room items fixed (During hypercare week)': '3', 'Date': '2/5/2022', 'Functional Ench': '41', 'Hotfix Prod\nDeploy Date': '14-Feb-22', 'Number of Hotfixes': '3', 'PDBs': '5', 'Post HCL Exemption': '3', 'Release': '15.1.0', 'Total Changes Deployed': '46'}
]

请问为何我的代码会在列表中生成重复值,该如何修复?


为什么会出现重复值?

这是Python里很常见的引用陷阱:你在循环开始前只创建了一个字典对象a,之后每处理一行数据,只是修改这个字典里的键值对,再把这个字典的引用添加到列表obj中。列表里的所有元素其实都指向同一个字典,所以最后列表里的所有字典都会显示最后一次循环修改后的内容,看起来就全是重复值了。

修复方案

核心解决思路是:每处理一行数据就创建一个全新的字典,而不是反复修改同一个字典。同时可以优化代码逻辑让它更简洁易读:

修复后的代码

import json
import pprint

fp = open('C:\\Users\\ADMIN\\Desktop\\test\\asd-Copy.txt')
data = json.load(fp)
obj = []

# 先提取表头字段,避免重复索引查询
headers = [col['v'] for col in data['rows'][0]]

# 跳过表头行,遍历所有数据行
for row in data['rows'][1:]:
    row_dict = {}  # 每行创建一个新字典,避免引用重复
    for idx, col in enumerate(row):
        current_header = headers[idx]
        if not col:
            # 空字典直接存入
            row_dict[current_header] = {}
        elif len(col) == 1:
            row_dict[current_header] = col['v']
        elif len(col) == 2:
            row_dict[current_header] = col['fv']
    obj.append(row_dict)

pprint.pprint(obj)

代码说明

  1. 移到循环内创建字典:把row_dict = {}放在每行数据的循环里,确保每行对应一个独立的字典对象,彻底解决引用重复问题。
  2. 表头提取优化:先把表头字段提取成列表headers,后续循环直接通过索引获取,代码更简洁高效。
  3. 条件判断优化:用elif替代多个if,避免同一元素被多次判断修改;空字典的判断用not col更直观。

这样修改后,列表里的每个字典都是独立的对象,输出结果会和你期望的完全一致。


内容的提问来源于stack exchange,提问作者Ranjan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 16:47:27