如何读取JSON文件中最新追加的数据?实现tail -f式读取
实现类似
tail -f的JSON文件实时追踪 嘿,我明白你想实现的是像tail -f那样实时读取JSON文件里新增的内容,而且你已经搞定了普通文本文件的情况——咱们来看看怎么适配JSON场景吧!
首先得明确你的JSON文件是哪种更新方式,不同的写入逻辑对应的读取方案差异很大:
情况1:JSON Lines格式(每行一个独立JSON对象)
如果你的更新脚本是每次追加一行完整的JSON对象(比如写入时用json.dumps(new_data) + '\n'),那这种场景和普通文本文件几乎一样,只是多了一步JSON解析。你可以基于你现有的代码修改:
import json import os import time class JsonLineTailer: def __init__(self, fileName): self.fileName = fileName # 打开文件并直接跳到末尾 self.file = open(self.fileName, 'r') self.file.seek(0, os.SEEK_END) def follow(self): while True: line = self.file.readline() if not line: print("No new data, waiting...") time.sleep(1) continue # 尝试解析JSON行,处理可能的格式错误 try: new_data = json.loads(line.strip()) print("New JSON entry:", new_data) except json.JSONDecodeError as e: print(f"Failed to parse line: {line.strip()}, error: {str(e)}") # 用法示例 if __name__ == "__main__": tailer = JsonLineTailer("your_updating.json") tailer.follow()
这种方案轻量高效,完全贴合tail -f的逻辑,是实时追踪JSON数据的最优选择。
情况2:往单个JSON数组里追加元素
如果你的JSON文件是一个数组结构(比如初始内容是[{"key": "value1"}],更新后变成[{"key": "value1"}, {"key": "value2"}]),直接用文本方式读新增字节会拿到不完整的JSON片段(比如可能读到, {"key": "value2"}]的一部分)。这时候可以用「定期重读文件+对比元素数量」的方案:
import json import time class JsonArrayTailer: def __init__(self, fileName): self.fileName = fileName self.last_element_count = 0 # 记录上次读取到的元素总数 def follow(self): while True: try: with open(self.fileName, 'r') as f: json_data = json.load(f) if not isinstance(json_data, list): print("JSON file is not an array, exiting...") break # 提取新增的元素 new_elements = json_data[self.last_element_count:] if new_elements: print("New elements added:", new_elements) self.last_element_count = len(json_data) except json.JSONDecodeError as e: print(f"Failed to parse JSON file: {str(e)}") except FileNotFoundError: print("Target file not found, waiting...") # 每隔1秒检查一次,可根据需求调整间隔时间 time.sleep(1) # 用法示例 if __name__ == "__main__": tailer = JsonArrayTailer("your_array.json") tailer.follow()
这个方案虽然会定期重读整个文件,但胜在逻辑简单,能保证JSON解析的正确性。如果你的JSON文件特别大,可以考虑优化成记录上次的文件偏移量,结合JSON结构定位新增内容,但复杂度会高很多。
你原来的文本文件tail -f代码大概是这样的:
self.fileName = fileName self.file = open(self.fileName, 'r') self.st_results = os.stat(fileName) self.st_size = self.st_results[6] self.file.seek(self.st_size) while 1: where = self.file.tell() line = self.file.readline() if not line: print "No line wait..."
最后给个小建议:如果可以的话,优先用JSON Lines格式存储更新的JSON数据,不管是写入还是读取都会更简单高效~
内容的提问来源于stack exchange,提问作者Aviral Srivastava
相关产品推荐
相关产品推荐

