You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取以换行分隔、包含多行JSON片段的类JSON文件?

处理多独立JSON对象文件的解决方案

问题场景

存在一个由日志工具追加生成的文件,文件内包含多个独立的JSON对象(每个对象可能是单行或多行JSON),整体并非合法JSON,但每个单独的JSON片段是有效的。示例文件内容如下:

{"date": "2022-11-29", "runs": [{"23597": 821260}, {"23617": 821699}]}
{"date": "2022-11-30", "runs": [{"23597": 821269}, {"23617": 8213534}]}

使用常规的json.load()读取会触发报错:

with open('run_log.json','r') as file:
    d = json.load(file)
    print(d)

报错信息:

JSONDecodeError: Extra data: line 3 column 1 (char 89)

解决方案

1. 读取所有条目

场景1:每个JSON对象占一行

逐行读取并解析每行的JSON:

import json

all_entries = []
with open('run_log.json', 'r') as file:
    for line in file:
        cleaned_line = line.strip()
        if cleaned_line:  # 跳过空行
            entry = json.loads(cleaned_line)
            all_entries.append(entry)

# 输出所有条目
print(all_entries)

场景2:JSON对象跨多行

使用json.JSONDecoder的raw_decode方法分段解析:

import json

def load_multiple_json(file_path):
    decoder = json.JSONDecoder()
    all_entries = []
    with open(file_path, 'r') as file:
        full_data = file.read()
        current_pos = 0
        total_length = len(full_data)
        
        while current_pos < total_length:
            try:
                obj, current_pos = decoder.raw_decode(full_data, current_pos)
                all_entries.append(obj)
                # 跳过解析后的空白字符
                while current_pos < total_length and full_data[current_pos].isspace():
                    current_pos += 1
            except json.JSONDecodeError:
                # 遇到无法解析的内容时终止,可根据需求调整错误处理逻辑
                break
    return all_entries

# 加载所有条目
all_entries = load_multiple_json('run_log.json')
print(all_entries)

2. 获取指定日期的runs列表

基于上述读取逻辑,直接过滤目标日期的条目:

单行JSON场景

import json

target_date = "2022-11-30"
target_runs = None

with open('run_log.json', 'r') as file:
    for line in file:
        cleaned_line = line.strip()
        if not cleaned_line:
            continue
        entry = json.loads(cleaned_line)
        if entry.get('date') == target_date:
            target_runs = entry.get('runs')
            break  # 找到目标后停止遍历

if target_runs:
    print(f"日期{target_date}的runs列表:{target_runs}")
else:
    print(f"未找到日期{target_date}的对应条目")

多行JSON场景

import json

def get_runs_by_date(file_path, target_date):
    decoder = json.JSONDecoder()
    with open(file_path, 'r') as file:
        full_data = file.read()
        current_pos = 0
        total_length = len(full_data)
        
        while current_pos < total_length:
            try:
                obj, current_pos = decoder.raw_decode(full_data, current_pos)
                if obj.get('date') == target_date:
                    return obj.get('runs')
                # 跳过空白字符
                while current_pos < total_length and full_data[current_pos].isspace():
                    current_pos += 1
            except json.JSONDecodeError:
                break
    return None

target_date = "2022-11-30"
target_runs = get_runs_by_date('run_log.json', target_date)

if target_runs:
    print(f"日期{target_date}的runs列表:{target_runs}")
else:
    print(f"未找到日期{target_date}的对应条目")

内容的提问来源于stack exchange,提问作者Dolliy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 10:25:36