You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历JSON文件合并统计userID时JSONDecodeError报错如何解决

问题描述

我是编程新手,恳请各位前辈耐心解答我的问题。
我需要遍历多个JSON文件,将它们格式化为数组对象后合并存储到一个新的JSON文件中,最终准确统计所有数据中唯一userID的数量。
目前我编写的代码如下:

for root, subdirs, files in os.walk("./"):
    for file in files:
        if file.endswith('.json'):
            to_queue = []
            with open(file, "r+") as f:
                print(file)
                old = f.read()
                f.seek(0)  # rewind
                # save to the old string after replace
                new = old.replace('}{', '},{')
                f.write(new)
                tmps = '[' + str(new) + ']'
                json_string = json.loads(tmps)
                for key in json_string:
                    to_queue.append(key)
                f.close
            with open('update.json', 'a') as file:
                json.dump(json_string, file, indent=2)
            with open('update.json') as f:
                data = json.load(f)
                users = set(item.get('userID') for item in data)
                print(len(users))
                # print(users

代码逻辑是遍历所有JSON文件,格式化后写入update.json,再读取该文件统计其中包含的唯一userID数量。但运行代码时报如下错误:

Traceback (most recent call last):
  File "format.py", line 26, in <module>
    data = json.load(f)
  File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 293, in load
    return loads(fp.read(),
  File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 357, in loads
    return _default_decoder.decode(s)
  File "/Users/user/opt/anaconda3/lib/python3.8/json/decoder.py", line 340, in decode
    raise JSONDecodeError("Extra data", s, end)
json.decoder.JSONDecodeError: Extra data: line 17821 column 2 (char 501079)

我后续参考他人建议调整了代码,修改后的代码如下:

for root, subdirs, files in os.walk("./"):
    for file in files:
        if file.endswith('.json'):
            to_queue = []
            newdictionary = {}
            with open(file, "r+") as f:
                print(file)
                old = f.read()
                f.seek(0)  # rewind
                # save to the old string after replace
                new = old.replace('}{', '},{')
                f.write(new)
                tmps = '[' + str(new) + ']'
                json_string = json.loads(tmps)
                for key in json_string:
                    to_queue.append(key)
                    newdictionary.update(key)
                f.close
            for key in newdictionary:
                with open('update.json', 'a') as file:
                    json.dump(key, file, indent=2)
            with open('update.json') as f:
                data = json.load(f)
                users = set(item.get('userID') for item in data)
                print(len(users))

修改后仍然报错,报错信息如下:

Traceback (most recent call last):
  File "format.py", line 40, in <module>
    data = json.load(f)
  File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 293, in load
    return loads(fp.read(),
  File "/Users/user/opt/anaconda3/lib/python3.8/json/__init__.py", line 357, in loads
    return _default_decoder.decode(s)
  File "/Users/user/opt/anaconda3/lib/python3.8/json/decoder.py", line 340, in decode
    raise JSONDecodeError("Extra data", s, end)
json.decoder.JSONDecodeError: Extra data: line 1 column 194 (char 193)

我的JSON源文件原始内容分为两种格式,第一种是行内连续JSON对象格式:

{"@timestamp":"2021-07-30T20:28:25.769Z","name":"","deviceAction":""},{"@timestamp":"2021-07-30T20:29:10.812Z","name":"","deviceAction":""}

第二种是带换行缩进的多JSON对象格式,部分对象包含额外字段:

{
    "@timestamp": "",
    "userID": "",
    "destinationUserName": "",
    "message": "",
    "name": ""
  },
  {
    "@timestamp": "",
    "userID": "",
    "destinationUserName": "",
    "message": "",
    "name": ""
  },
  {
    "@timestamp": "",
    "userID": "",
    "destinationUserName": "",
    "message": "",
    "name": ""
  },
  {
    "@timestamp": "",
    "userID": "",
    "name": "",
    "sourceUserName": "",
    "deviceAction": ""
  }

我之前已经做了替换拼接转JSON数组的处理,还是出现上述报错,麻烦帮忙分析报错原因以及对应的解决方法,非常感谢。


问题解答

报错原因

  • 核心原因是update.json的内容不符合标准JSON格式要求:JSON只允许存在一个根元素,你使用追加模式a多次向update.json写入JSON内容,最终文件内会出现多个独立的JSON结构拼接的情况,解析时就会抛出Extra data错误。
  • 第一次代码的具体问题:每处理一个源JSON文件,就把解析得到的数组直接追加写入update.json,最终文件内容类似[{"a":1},{"b":2}][{"c":3},{"d":4}],存在多个根数组,无法解析。
  • 第二次修改后的代码问题更大:你错误遍历了字典的键写入文件,写入的根本不是合法的JSON对象结构,直接打乱了文件格式。
  • 额外的潜在问题:你使用r+模式修改原JSON文件时只调用了seek(0)没有调用truncate(),如果修改后的内容比原内容短,会残留旧内容,导致原文件损坏,而且修改源文件完全没有必要,你只需要读取内容做处理即可,不需要回写。

解决方法

直接在内存中累计所有数据,最后一次性写入结果文件即可,不需要中途反复读写update.json,修复后的代码如下:

import os
import json

# 初始化全局变量,累计所有数据和唯一userID
all_data = []
user_ids = set()

for root, subdirs, files in os.walk("./"):
    for file in files:
        # 跳过结果文件,避免重复读取
        if file == 'update.json' or not file.endswith('.json'):
            continue
        file_path = os.path.join(root, file)
        print(f"处理文件:{file_path}")
        with open(file_path, "r", encoding="utf-8") as f:
            old_content = f.read()
        # 处理连续JSON对象,转为数组
        formatted_content = f"[{old_content.replace('}{', '},{')}]"
        try:
            json_array = json.loads(formatted_content)
        except Exception as e:
            print(f"文件{file_path}解析失败,错误:{e},跳过该文件")
            continue
        # 累计数据和userID
        for item in json_array:
            all_data.append(item)
            user_id = item.get('userID')
            if user_id:
                user_ids.add(user_id)

# 一次性写入合并后的结果
with open('update.json', 'w', encoding="utf-8") as f:
    json.dump(all_data, f, indent=2, ensure_ascii=False)

# 输出统计结果
print(f"所有JSON文件合并完成,共{len(all_data)}条数据")
print(f"唯一userID的数量为:{len(user_ids)}")

代码说明

  • 全程只读取源JSON文件,不修改源文件内容,避免损坏原始数据
  • 跳过update.json本身,避免把结果文件当成源文件处理
  • 增加了异常捕获,单个文件解析失败不会导致整个程序崩溃
  • 所有数据在内存中累计,最后一次性写入结果文件,保证update.json是合法的单根JSON数组
  • 直接在遍历数据时累计userID,不需要后续再读取一次结果文件,效率更高

内容的提问来源于stack exchange,提问作者Nayden Van

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 00:24:03