JSON转CSV格式异常:Python实现嵌套对象对应单元格上移
JSON转CSV格式修复:单元格上移实现方案
问题说明
将包含嵌套对象的JSON列表转换为CSV时出现格式问题:输出存在冗余空行、列错位。JSON对象除可选的Dict2外共享相同主键,Dict1可为带嵌套子字典的字典或字符串,目标是生成规范的全数据CSV。希望尽量不重写现有递归代码,通过Python实现指定单元格上移修复格式。
原JSON结构
[ { "Name": "NameOfObject1", "Number": "objectNumber1", "Dict1": { "Key1": "true", "Key2": "false", "Key3": { "key1InK3": "Hello", "key2InK3": "There" } }, "BooleanValue": false }, { "Name": "NameOfObject2", "Number": "objectNumber2", "Dict1": { "Key1": "true", "Key2": "false", "Key3": "true" }, "BooleanValue": false, "Dict2": { "otherKey1": "this ", "otherKey2": "is ", "otherKey3": "a ", "otherKey4": "test ", "otherKey5": "haha ", "otherKey6": "i ", "otherKey7": "like ", "otherKey8": "strawberry " } }, { "Name": "NameOfObject3", "Number": "objectNumber", "Dict1": "this is a string", "BooleanValue": false } ]
当前错误CSV输出
Name Number Dic1 Dic1 Values Dic1 Values Of dictionaries Inside Boolean Value Dic2 Dic2 Values NameOfObject1 objectNumber1 FALSE NameOfObject1 objectNumber1 Key1 TRUE FALSE NameOfObject1 objectNumber1 Key2 FALSE NameOfObject1 objectNumber1 Key2 key1InK2 FALSE NameOfObject1 objectNumber1 Key2 key1InK2 Hello FALSE NameOfObject1 objectNumber1 Key2 key2InK2 FALSE NameOfObject1 objectNumber1 Key2 key2InK2 There FALSE NameOfObject2 objectNumber2 TRUE NameOfObject1 objectNumber1 Key3 FALSE NameOfObject1 objectNumber1 Key3 key1InK3 FALSE NameOfObject1 objectNumber1 Key3 key1InK3 Good FALSE NameOfObject1 objectNumber1 Key3 key2InK3 FALSE NameOfObject1 objectNumber1 Key3 key2InK3 Morining FALSE NameOfObject2 objectNumber2 Key1 TRUE TRUE NameOfObject2 objectNumber2 Key2 FALSE TRUE NameOfObject2 objectNumber2 Key3 TRUE TRUE NameOfObject2 objectNumber2 TRUE otherKey1 This NameOfObject2 objectNumber2 TRUE otherKey2 is NameOfObject2 objectNumber2 TRUE otherKey3 a NameOfObject2 objectNumber2 TRUE otherKey4 test NameOfObject2 objectNumber2 TRUE otherKey5 haha NameOfObject2 objectNumber2 TRUE otherKey6 i NameOfObject2 objectNumber2 TRUE otherKey7 like NameOfObject2 objectNumber2 TRUE otherKey8 strawberry NameOfObject3 objectNumber3 This is a String FALSE
期望CSV输出
Name Number Dic1 Dic1 Values Dic1 Values Of dictionaries Inside Boolean Value Dic2 Dic2 Values NameOfObject1 objectNumber1 FALSE NameOfObject1 objectNumber1 Key1 TRUE FALSE NameOfObject1 objectNumber1 Key2 key1InK2 Hello FALSE NameOfObject1 objectNumber1 Key2 key2InK2 There FALSE NameOfObject2 objectNumber2 TRUE NameOfObject1 objectNumber1 Key3 key1InK3 Good FALSE NameOfObject1 objectNumber1 Key3 key2InK3 Morining FALSE NameOfObject2 objectNumber2 TRUE TRUE NameOfObject2 objectNumber2 Key1 TRUE TRUE otherKey1 this NameOfObject2 objectNumber2 Key2 FALSE TRUE otherKey2 is NameOfObject2 objectNumber2 Key3 TRUE TRUE otherKey3 a NameOfObject2 objectNumber2 TRUE otherKey4 test NameOfObject2 objectNumber2 TRUE otherKey5 haha NameOfObject2 objectNumber2 TRUE otherKey6 i NameOfObject2 objectNumber2 TRUE otherKey7 like NameOfObject2 objectNumber2 TRUE otherKey8 strawberry NameOfObject2 objectNumber2 TRUE otherKey i NameOfObject2 objectNumber2 TRUE otherKey7 like NameOfObject2 objectNumber2 TRUE otherKey8 strawberry NameOfObject3 objectNumber3 This is a String FALSE
现有核心代码逻辑
ListOfLines = [] # 遍历JSON对象的键 for key in dictionary_keys: if isinstance(value, dict): # 如果是字典,追加新行并递归调用处理函数 ListOfLines.append(new_line) recursive_function(value) else: # 非字典类型,将键值添加到当前行 add_key_value_to_current_line(key, value) # 格式化所有行并清理列表,最后写入输出文件 format_lines_and_write_to_file(ListOfLines)
解决方案:后处理行列表实现单元格上移
不修改现有递归逻辑,直接对生成的ListOfLines做后处理,完成空行过滤和单元格上移填充:
具体实现代码
def clean_and_shift_lines(lines): # 1. 过滤空行(只保留非空且非全空格的行) cleaned = [line.strip().split('\t') for line in lines if line.strip()] if not cleaned: return [] header = cleaned[0] col_count = len(header) processed = [header] # 2. 按Name和Number分组,同一对象的行归为一组 groups = {} for row in cleaned[1:]: # 补全行的列数,避免索引越界 while len(row) < col_count: row.append('') name, number = row[0], row[1] key = (name, number) if key not in groups: groups[key] = [] groups[key].append(row) # 3. 处理每组,实现单元格上移填充 for group in groups.values(): current_row = group[0].copy() for row in group[1:]: # 用当前行的非空值填充前一行的空列 for idx in range(col_count): if not current_row[idx] and row[idx]: current_row[idx] = row[idx] # 判断当前行是否有新内容,有则保留当前填充行,重置current_row has_new = any(row[idx] and row[idx] != current_row[idx] for idx in range(col_count)) if has_new: processed.append(current_row) current_row = row.copy() # 添加每组最后一行 processed.append(current_row) # 转回制表符分隔的字符串 return ['\t'.join(row) for row in processed] # 使用示例:将现有代码生成的ListOfLines传入处理 # fixed_lines = clean_and_shift_lines(ListOfLines) # with open('output.csv', 'w', encoding='utf-8') as f: # f.write('\n'.join(fixed_lines))
代码说明
- 先过滤空行,避免冗余数据干扰
- 按
Name和Number分组,确保同一对象的行一起处理 - 组内遍历行,把后续行的非空值往前填充到前面行的空列,同时移除仅包含重复空值的冗余行
- 最后将处理后的行转回CSV格式字符串,写入文件即可
内容的提问来源于stack exchange,提问作者ahmad
相关产品推荐
相关产品推荐

