You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分包含多个JSON数据的字符串文件

拆分连续拼接的JSON文件解决方案

问题场景

处理一个包含1-2个连续有效JSON字符串的文件,文件内容为纯JSON直接拼接(无分隔符、无类结构),需要拆分出各自独立的JSON字符串。示例文件内容如下:

{
    "notes_report": {
        "Version": "1.0",
        "SubscriberId": 123456,
        "ItemId": 98745632,
        "Allocations": [
            {
                "AllocatedAdjustmentId": "98745632",
                "AllocatedAdjustmentNumber": "DN00412345698",
                "Date": "2022-10-11T07:29:00.000Z",
                "AmountPaid": 3.0,
                "Balance": 2.0,
                "OriginalBalance": -5.0
            }
        ],
        "ReferenceId": "CN00432165498",
        "NoteType": 1,
        "AccountNumber": "ACCQ1023654789",
        "Amount": -10.0,
        "RemainingBalance": -7.0,
        "Currency": "USD",
        "CreatedDate": "2022-10-11T07:29:00.000Z"
    },
    "BusinessUnitId": 255,
    "EventNotificationType": 2139,
    "EventDate": "2022-10-13T16:57:33.873Z"
}
{
    "notes_report": {
        "Version": "1.0",
        "SubscriberId": 741085,
        "ItemId": 32014587,
        "Allocations": [
            {
                "AllocatedAdjustmentId": "98741023",
                "AllocatedAdjustmentNumber": "CN00432165498",
                "Date": "2022-10-11T07:29:00.000Z",
                "AmountPaid": -3.0,
                "Balance": 10.0,
                "OriginalBalance": 10.0
            }
        ],
        "ReferenceId": "DN00400146452",
        "NoteType": 2,
        "AccountNumber": "ACCQ1023654789",
        "Amount": 5.0,
        "RemainingBalance": 5.0,
        "Currency": "USD",
        "CreatedDate": "2022-10-11T08:50:00.000Z"
    },
    "BusinessUnitId": 255,
    "EventNotificationType": 2139,
    "EventDate": "2022-10-13T16:57:33.896Z"
}

解决思路

核心逻辑是通过计数大括号的开闭状态确定JSON边界:每个完整JSON以{开头,以匹配的}结尾,当括号计数回到0时,即为一个JSON的结束位置。由于文件最多包含2个JSON,找到边界后直接拆分即可。

代码实现(Python)

def split_concat_json(file_path):
    # 读取文件内容
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read().strip()
    
    json_list = []
    bracket_count = 0
    start_idx = 0
    
    # 遍历字符,追踪大括号计数
    for idx, char in enumerate(content):
        if char == '{':
            bracket_count += 1
        elif char == '}':
            bracket_count -= 1
            # 计数归0时,截取完整JSON
            if bracket_count == 0:
                json_str = content[start_idx:idx+1].strip()
                json_list.append(json_str)
                start_idx = idx + 1
                # 最多两个JSON,提前终止遍历
                if len(json_list) == 2:
                    break
    return json_list

# 使用示例
if __name__ == "__main__":
    # 替换为你的文件路径
    result = split_concat_json("your_json_file.txt")
    # 输出并保存拆分后的JSON
    for index, json_str in enumerate(result, 1):
        print(f"拆分出的第{index}个JSON:\n{json_str}\n")
        with open(f"split_json_{index}.json", 'w', encoding='utf-8') as f:
            f.write(json_str)

手动拆分方法

如果仅处理单个文件,可直接定位第一个JSON的最后一个}:观察示例内容,第一个JSON的结束位置是第一块内容的最后一行},直接在此处分割即可得到两个独立的JSON字符串。

内容的提问来源于stack exchange,提问作者BMD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 17:40:52