Python中从文件JSON数组获取新增追加元素的最优方法
监控JSON数组文件获取新元素的实用方法
方法1:按文件大小增量读取
核心思路是记录每次读取后的文件大小,下次只读取新增的部分,再解析成JSON元素,适合持续往数组末尾追加内容的场景。
步骤很简单:
- 第一次读完整文件,解析成JSON数组,记下文件大小和数组长度
- 定时检查文件大小,发现变大就读新增的字节
- 处理新增内容的格式(比如去掉末尾的逗号、补全JSON结构),提取新元素
代码示例:
import json import time import os def monitor_json_file(file_path, check_interval=2): # 初始化:读取全量数据,记录初始状态 last_file_size = os.path.getsize(file_path) with open(file_path, 'r', encoding='utf-8') as f: full_data = json.load(f) last_array_length = len(full_data) while True: current_size = os.path.getsize(file_path) if current_size > last_file_size: # 读取新增的内容 with open(file_path, 'r', encoding='utf-8') as f: f.seek(last_file_size) new_raw_content = f.read() # 清理格式:去掉末尾可能的逗号和未闭合的方括号 cleaned_content = new_raw_content.rstrip().rstrip(',').rstrip(']') if not cleaned_content: last_file_size = current_size time.sleep(check_interval) continue # 拆分并解析新元素 try: if ',' in cleaned_content: new_elements = [json.loads(elem.strip()) for elem in cleaned_content.split(',')] else: new_elements = [json.loads(cleaned_content.strip())] # 输出新元素,这里可以替换成你的业务逻辑 for elem in new_elements: print("抓到新元素:", elem) last_array_length += len(new_elements) last_file_size = current_size except json.JSONDecodeError: # 写入过程中文件可能不完整,跳过这次 pass time.sleep(check_interval) # 调用示例,替换成你的JSON文件路径 monitor_json_file('your_data.json')
方法2:用文件事件监控(watchdog库)
利用系统的文件修改事件触发读取,不用定时轮询,更高效,适合可能重写整个文件但数组是追加元素的场景。
先安装依赖:pip install watchdog
代码示例:
import json from watchdog.observers import Observer from watchdog.events import FileSystemEventHandler class JSONMonitorHandler(FileSystemEventHandler): def __init__(self, target_file): self.target_file = target_file # 初始化读取全量数据 with open(target_file, 'r', encoding='utf-8') as f: self.last_data = json.load(f) self.last_length = len(self.last_data) def on_modified(self, event): # 只处理目标文件的修改事件,忽略目录 if not event.is_directory and event.src_path == self.target_file: try: with open(self.target_file, 'r', encoding='utf-8') as f: current_data = json.load(f) # 对比数组长度,提取新元素 if len(current_data) > self.last_length: new_items = current_data[self.last_length:] for item in new_items: print("新增元素:", item) self.last_length = len(current_data) self.last_data = current_data except json.JSONDecodeError: # 文件正在写入,跳过本次读取 pass def start_json_monitor(file_path): handler = JSONMonitorHandler(file_path) observer = Observer() # 监控文件所在目录 observer.schedule(handler, path='.', recursive=False) observer.start() try: # 保持进程运行 while True: pass except KeyboardInterrupt: observer.stop() observer.join() # 调用示例 start_json_monitor('your_data.json')
方法3:改成逐行写入的日志格式(最省心)
如果能控制写入方,别再维护大数组了,改成每行一个JSON对象,读取的时候直接像读日志一样 tail 新增行,实现最简单,性能也最好。
调整后的文件格式:
{"key1":"value1","key2":"value2"} {"key1":"value3","key2":"value4"}
代码示例:
import json import time def tail_json_lines(file_path, check_interval=2): with open(file_path, 'r', encoding='utf-8') as f: # 跳到文件末尾 f.seek(0, 2) while True: line = f.readline() if line: try: item = json.loads(line.strip()) print("新增元素:", item) except json.JSONDecodeError: # 跳过无效行 continue time.sleep(check_interval) # 调用示例 tail_json_lines('your_data_lines.json')
内容的提问来源于stack exchange,提问作者novice
相关产品推荐
相关产品推荐

