You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取JSON文件中最新追加的数据?实现tail -f式读取

实现类似tail -f的JSON文件实时追踪

嘿,我明白你想实现的是像tail -f那样实时读取JSON文件里新增的内容,而且你已经搞定了普通文本文件的情况——咱们来看看怎么适配JSON场景吧!

首先得明确你的JSON文件是哪种更新方式,不同的写入逻辑对应的读取方案差异很大:

情况1:JSON Lines格式(每行一个独立JSON对象)

如果你的更新脚本是每次追加一行完整的JSON对象(比如写入时用json.dumps(new_data) + '\n'),那这种场景和普通文本文件几乎一样,只是多了一步JSON解析。你可以基于你现有的代码修改:

import json
import os
import time

class JsonLineTailer:
    def __init__(self, fileName):
        self.fileName = fileName
        # 打开文件并直接跳到末尾
        self.file = open(self.fileName, 'r')
        self.file.seek(0, os.SEEK_END)
        
    def follow(self):
        while True:
            line = self.file.readline()
            if not line:
                print("No new data, waiting...")
                time.sleep(1)
                continue
            # 尝试解析JSON行,处理可能的格式错误
            try:
                new_data = json.loads(line.strip())
                print("New JSON entry:", new_data)
            except json.JSONDecodeError as e:
                print(f"Failed to parse line: {line.strip()}, error: {str(e)}")

# 用法示例
if __name__ == "__main__":
    tailer = JsonLineTailer("your_updating.json")
    tailer.follow()

这种方案轻量高效,完全贴合tail -f的逻辑,是实时追踪JSON数据的最优选择。

情况2:往单个JSON数组里追加元素

如果你的JSON文件是一个数组结构(比如初始内容是[{"key": "value1"}],更新后变成[{"key": "value1"}, {"key": "value2"}]),直接用文本方式读新增字节会拿到不完整的JSON片段(比如可能读到, {"key": "value2"}]的一部分)。这时候可以用「定期重读文件+对比元素数量」的方案:

import json
import time

class JsonArrayTailer:
    def __init__(self, fileName):
        self.fileName = fileName
        self.last_element_count = 0  # 记录上次读取到的元素总数
        
    def follow(self):
        while True:
            try:
                with open(self.fileName, 'r') as f:
                    json_data = json.load(f)
                    if not isinstance(json_data, list):
                        print("JSON file is not an array, exiting...")
                        break
                    # 提取新增的元素
                    new_elements = json_data[self.last_element_count:]
                    if new_elements:
                        print("New elements added:", new_elements)
                        self.last_element_count = len(json_data)
            except json.JSONDecodeError as e:
                print(f"Failed to parse JSON file: {str(e)}")
            except FileNotFoundError:
                print("Target file not found, waiting...")
            # 每隔1秒检查一次,可根据需求调整间隔时间
            time.sleep(1)

# 用法示例
if __name__ == "__main__":
    tailer = JsonArrayTailer("your_array.json")
    tailer.follow()

这个方案虽然会定期重读整个文件,但胜在逻辑简单,能保证JSON解析的正确性。如果你的JSON文件特别大,可以考虑优化成记录上次的文件偏移量,结合JSON结构定位新增内容,但复杂度会高很多。


你原来的文本文件tail -f代码大概是这样的:

self.fileName = fileName 
self.file = open(self.fileName, 'r') 
self.st_results = os.stat(fileName) 
self.st_size = self.st_results[6] 
self.file.seek(self.st_size) 
while 1: 
    where = self.file.tell() 
    line = self.file.readline() 
    if not line: 
        print "No line wait..."

最后给个小建议:如果可以的话,优先用JSON Lines格式存储更新的JSON数据,不管是写入还是读取都会更简单高效~

内容的提问来源于stack exchange,提问作者Aviral Srivastava

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:26:40