Python中如何更快且低体积存储图像numpy数组、数据及时间戳?
更优的numpy图像与数据存储方案?
我需要在Python中将numpy数组(图像)、数据和时间戳存储到文件中,后续能浏览图像并查看对应的数据。目前用JSON文件存储,但写入耗时久,一张640x480的图像加一组数据就占10MB体积,以下是我的实现代码:
import json as json import base64 import io import sys import cv2 import numpy as np import os from numpy import asarray from json import JSONEncoder path = r'C:/Videotool/Data' name = 'testfile' cam = cv2.VideoCapture(0) result, image = cam.read() if result: print(f'image:{image}') else: print('no image') def json_write_to_file(path, name, data): file_exists = os.path.exists(name+'.json') file = path + '/' + name + '.json' readed_data = [] data_list = [] sum_dict = [] if file_exists: print(f'file {name}.json exist\n') with open(file) as fp: readed_data = json.load(fp) print(f'readed_data:{readed_data}\n') print(f'data:{data}\n') readed_data.append(data) sum_dict = readed_data print(f'sum_dict:{sum_dict}\n') with open (file, 'w') as fp: json.dump(sum_dict, fp, indent=2) print(f'length of readed_data: {len(readed_data)}\n') #fp.write('\n') else: sum_dict.append(data) print(f'file {name}.json not exists\n') with open (file, 'w') as fp: json.dump(sum_dict, fp, indent=2) print(f'Number of Dimensions of image:{image.ndim}') print(f'Shape of image:{image.shape}') print(f'Size of image:{image.size}') low_image_array = image.reshape(-1) ori_array = np.zeros((2,3,4)) print(f'Number of Dimensions of ori_array:{ori_array.ndim}') print(f'Shape of ori_array:{ori_array.shape}') print(f'Size of ori_array:{ori_array.size}') print(f'ori_array:{ori_array}') low_1_array = ori_array.reshape(-1) print(f'Number of Dimensions of low_1_array:{low_1_array.ndim}') print(f'Shape of low_1_array:{low_1_array.shape}') print(f'Size of low_1_array:{low_1_array.size}') print(f'low_1_array:{low_1_array}') data_set = {} data_set['Stream_1'] = low_1_array.tolist() data_set['Stream_2'] = low_image_array.tolist() print(f'data_set: {data_set}') json_write_to_file(path, name, data_set)
推荐的存储方案
1. numpy原生.npy/.npz格式
专为numpy数组设计,存储效率和读写速度远胜JSON,还能完整保留数组的形状、 dtype等原生信息。
- 实现代码:
import numpy as np import cv2 import time # 获取图像、数据与时间戳 cam = cv2.VideoCapture(0) result, image = cam.read() data_array = np.zeros((2,3,4)) timestamp = time.time() # 打包存储(自动压缩) np.savez("data.npz", image=image, stream1=data_array, timestamp=timestamp) # 读取还原 loaded_data = np.load("data.npz") loaded_image = loaded_data["image"] loaded_stream1 = loaded_data["stream1"] loaded_timestamp = loaded_data["timestamp"]
- 优势:体积仅为JSON的1/5~1/10,读写速度提升数倍,无需手动转换数组格式。
2. HDF5格式(h5py库)
适合长期归档、增量写入的大规模数据场景,支持分层管理数据和压缩存储。
- 先安装依赖:
pip install h5py - 实现代码:
import h5py import cv2 import time cam = cv2.VideoCapture(0) result, image = cam.read() data_array = np.zeros((2,3,4)) timestamp = time.time() # 增量写入(文件存在则追加记录) with h5py.File("data.h5", "a") as f: # 按时间戳创建独立数据组 group = f.create_group(f"record_{int(timestamp)}") group.create_dataset("image", data=image, compression="gzip") group.create_dataset("stream1", data=data_array) group.attrs["timestamp"] = timestamp # 读取数据 with h5py.File("data.h5", "r") as f: for group_name in f: group = f[group_name] loaded_image = group["image"][()] loaded_stream1 = group["stream1"][()] loaded_timestamp = group.attrs["timestamp"]
- 优势:支持无损压缩,可高效管理数千条以上的记录,适合专业数据存储场景。
3. 图像单独存储+元数据轻量记录
适合需要直接用图像浏览器查看图像的场景,把图像存为JPG/PNG,元数据用CSV或SQLite记录。
- 实现代码:
import cv2 import csv import time import os import numpy as np cam = cv2.VideoCapture(0) result, image = cam.read() data_array = np.zeros((2,3,4)) timestamp = time.time() # 保存图像为JPG(体积远小于JSON存储) img_filename = f"image_{int(timestamp)}.jpg" cv2.imwrite(img_filename, image) # 元数据写入CSV data_row = [ timestamp, img_filename, ",".join(map(str, data_array.flatten())) ] file_exists = os.path.exists("metadata.csv") with open("metadata.csv", "a", newline="") as f: writer = csv.writer(f) if not file_exists: writer.writerow(["timestamp", "image_filename", "stream1_data"]) writer.writerow(data_row) # 读取还原 with open("metadata.csv", "r") as f: reader = csv.DictReader(f) for row in reader: ts = float(row["timestamp"]) img = cv2.imread(row["image_filename"]) stream1 = np.array(list(map(float, row["stream1_data"].split(",")))).reshape(2,3,4)
- 优势:图像可直接浏览,元数据文件体积极小,读写速度极快。
内容的提问来源于stack exchange,提问作者Chorum
相关产品推荐
相关产品推荐

