Python遍历JSON文件合并统计userID时JSONDecodeError报错如何解决
问题描述
我是编程新手,恳请各位前辈耐心解答我的问题。
我需要遍历多个JSON文件,将它们格式化为数组对象后合并存储到一个新的JSON文件中,最终准确统计所有数据中唯一userID的数量。
目前我编写的代码如下:
for root, subdirs, files in os.walk("./"): for file in files: if file.endswith('.json'): to_queue = [] with open(file, "r+") as f: print(file) old = f.read() f.seek(0) # rewind # save to the old string after replace new = old.replace('}{', '},{') f.write(new) tmps = '[' + str(new) + ']' json_string = json.loads(tmps) for key in json_string: to_queue.append(key) f.close with open('update.json', 'a') as file: json.dump(json_string, file, indent=2) with open('update.json') as f: data = json.load(f) users = set(item.get('userID') for item in data) print(len(users)) # print(users
代码逻辑是遍历所有JSON文件,格式化后写入update.json,再读取该文件统计其中包含的唯一userID数量。但运行代码时报如下错误:
Traceback (most recent call last): File "format.py", line 26, in <module> data = json.load(f) File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 293, in load return loads(fp.read(), File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 357, in loads return _default_decoder.decode(s) File "/Users/user/opt/anaconda3/lib/python3.8/json/decoder.py", line 340, in decode raise JSONDecodeError("Extra data", s, end) json.decoder.JSONDecodeError: Extra data: line 17821 column 2 (char 501079)
我后续参考他人建议调整了代码,修改后的代码如下:
for root, subdirs, files in os.walk("./"): for file in files: if file.endswith('.json'): to_queue = [] newdictionary = {} with open(file, "r+") as f: print(file) old = f.read() f.seek(0) # rewind # save to the old string after replace new = old.replace('}{', '},{') f.write(new) tmps = '[' + str(new) + ']' json_string = json.loads(tmps) for key in json_string: to_queue.append(key) newdictionary.update(key) f.close for key in newdictionary: with open('update.json', 'a') as file: json.dump(key, file, indent=2) with open('update.json') as f: data = json.load(f) users = set(item.get('userID') for item in data) print(len(users))
修改后仍然报错,报错信息如下:
Traceback (most recent call last): File "format.py", line 40, in <module> data = json.load(f) File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 293, in load return loads(fp.read(), File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 357, in loads return _default_decoder.decode(s) File "/Users/user/opt/anaconda3/lib/python3.8/json/decoder.py", line 340, in decode raise JSONDecodeError("Extra data", s, end) json.decoder.JSONDecodeError: Extra data: line 1 column 194 (char 193)
我的JSON源文件原始内容分为两种格式,第一种是行内连续JSON对象格式:
{"@timestamp":"2021-07-30T20:28:25.769Z","name":"","deviceAction":""},{"@timestamp":"2021-07-30T20:29:10.812Z","name":"","deviceAction":""}
第二种是带换行缩进的多JSON对象格式,部分对象包含额外字段:
{ "@timestamp": "", "userID": "", "destinationUserName": "", "message": "", "name": "" }, { "@timestamp": "", "userID": "", "destinationUserName": "", "message": "", "name": "" }, { "@timestamp": "", "userID": "", "destinationUserName": "", "message": "", "name": "" }, { "@timestamp": "", "userID": "", "name": "", "sourceUserName": "", "deviceAction": "" }
我之前已经做了替换拼接转JSON数组的处理,还是出现上述报错,麻烦帮忙分析报错原因以及对应的解决方法,非常感谢。
问题解答
报错原因
- 核心原因是
update.json的内容不符合标准JSON格式要求:JSON只允许存在一个根元素,你使用追加模式a多次向update.json写入JSON内容,最终文件内会出现多个独立的JSON结构拼接的情况,解析时就会抛出Extra data错误。 - 第一次代码的具体问题:每处理一个源JSON文件,就把解析得到的数组直接追加写入
update.json,最终文件内容类似[{"a":1},{"b":2}][{"c":3},{"d":4}],存在多个根数组,无法解析。 - 第二次修改后的代码问题更大:你错误遍历了字典的键写入文件,写入的根本不是合法的JSON对象结构,直接打乱了文件格式。
- 额外的潜在问题:你使用
r+模式修改原JSON文件时只调用了seek(0)没有调用truncate(),如果修改后的内容比原内容短,会残留旧内容,导致原文件损坏,而且修改源文件完全没有必要,你只需要读取内容做处理即可,不需要回写。
解决方法
直接在内存中累计所有数据,最后一次性写入结果文件即可,不需要中途反复读写update.json,修复后的代码如下:
import os import json # 初始化全局变量,累计所有数据和唯一userID all_data = [] user_ids = set() for root, subdirs, files in os.walk("./"): for file in files: # 跳过结果文件,避免重复读取 if file == 'update.json' or not file.endswith('.json'): continue file_path = os.path.join(root, file) print(f"处理文件:{file_path}") with open(file_path, "r", encoding="utf-8") as f: old_content = f.read() # 处理连续JSON对象,转为数组 formatted_content = f"[{old_content.replace('}{', '},{')}]" try: json_array = json.loads(formatted_content) except Exception as e: print(f"文件{file_path}解析失败,错误:{e},跳过该文件") continue # 累计数据和userID for item in json_array: all_data.append(item) user_id = item.get('userID') if user_id: user_ids.add(user_id) # 一次性写入合并后的结果 with open('update.json', 'w', encoding="utf-8") as f: json.dump(all_data, f, indent=2, ensure_ascii=False) # 输出统计结果 print(f"所有JSON文件合并完成,共{len(all_data)}条数据") print(f"唯一userID的数量为:{len(user_ids)}")
代码说明
- 全程只读取源JSON文件,不修改源文件内容,避免损坏原始数据
- 跳过
update.json本身,避免把结果文件当成源文件处理 - 增加了异常捕获,单个文件解析失败不会导致整个程序崩溃
- 所有数据在内存中累计,最后一次性写入结果文件,保证
update.json是合法的单根JSON数组 - 直接在遍历数据时累计userID,不需要后续再读取一次结果文件,效率更高
内容的提问来源于stack exchange,提问作者Nayden Van
相关产品推荐
相关产品推荐

