You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何更快且低体积存储图像numpy数组、数据及时间戳?

更优的numpy图像与数据存储方案?

我需要在Python中将numpy数组(图像)、数据和时间戳存储到文件中,后续能浏览图像并查看对应的数据。目前用JSON文件存储,但写入耗时久,一张640x480的图像加一组数据就占10MB体积,以下是我的实现代码:

import json as json
import base64
import io
import sys
import cv2

import numpy as np

import os

from numpy import asarray
from json import JSONEncoder

path = r'C:/Videotool/Data'
name = 'testfile'


cam = cv2.VideoCapture(0)
result, image = cam.read()

if result:
    print(f'image:{image}')
else:
    print('no image')


def json_write_to_file(path, name, data):
        file_exists = os.path.exists(name+'.json')
        file = path + '/' + name + '.json'
        readed_data = []
        data_list = []
        sum_dict = []
        if file_exists:
            print(f'file {name}.json exist\n')
            with open(file) as fp:
                readed_data = json.load(fp)
            print(f'readed_data:{readed_data}\n')
            
            print(f'data:{data}\n')

            readed_data.append(data)
            sum_dict = readed_data

            print(f'sum_dict:{sum_dict}\n')
            with open (file, 'w') as fp:
                json.dump(sum_dict, fp, indent=2)
                print(f'length of readed_data: {len(readed_data)}\n')
                #fp.write('\n')
        else:
            sum_dict.append(data)
            print(f'file {name}.json not exists\n')    
            with open (file, 'w') as fp:
                json.dump(sum_dict, fp, indent=2)


print(f'Number of Dimensions of image:{image.ndim}')
print(f'Shape of image:{image.shape}')
print(f'Size of image:{image.size}')

low_image_array = image.reshape(-1)

ori_array = np.zeros((2,3,4))

print(f'Number of Dimensions of ori_array:{ori_array.ndim}')
print(f'Shape of ori_array:{ori_array.shape}')
print(f'Size of ori_array:{ori_array.size}')

print(f'ori_array:{ori_array}')

low_1_array = ori_array.reshape(-1)


print(f'Number of Dimensions of low_1_array:{low_1_array.ndim}')
print(f'Shape of low_1_array:{low_1_array.shape}')
print(f'Size of low_1_array:{low_1_array.size}')

print(f'low_1_array:{low_1_array}')

data_set = {}
data_set['Stream_1']    = low_1_array.tolist()
data_set['Stream_2']    = low_image_array.tolist()

print(f'data_set: {data_set}')

json_write_to_file(path, name, data_set)

推荐的存储方案

1. numpy原生.npy/.npz格式

专为numpy数组设计,存储效率和读写速度远胜JSON,还能完整保留数组的形状、 dtype等原生信息。

  • 实现代码:
import numpy as np
import cv2
import time

# 获取图像、数据与时间戳
cam = cv2.VideoCapture(0)
result, image = cam.read()
data_array = np.zeros((2,3,4))
timestamp = time.time()

# 打包存储(自动压缩)
np.savez("data.npz", image=image, stream1=data_array, timestamp=timestamp)

# 读取还原
loaded_data = np.load("data.npz")
loaded_image = loaded_data["image"]
loaded_stream1 = loaded_data["stream1"]
loaded_timestamp = loaded_data["timestamp"]
  • 优势:体积仅为JSON的1/5~1/10,读写速度提升数倍,无需手动转换数组格式。

2. HDF5格式(h5py库)

适合长期归档、增量写入的大规模数据场景,支持分层管理数据和压缩存储。

  • 先安装依赖:pip install h5py
  • 实现代码:
import h5py
import cv2
import time

cam = cv2.VideoCapture(0)
result, image = cam.read()
data_array = np.zeros((2,3,4))
timestamp = time.time()

# 增量写入(文件存在则追加记录)
with h5py.File("data.h5", "a") as f:
    # 按时间戳创建独立数据组
    group = f.create_group(f"record_{int(timestamp)}")
    group.create_dataset("image", data=image, compression="gzip")
    group.create_dataset("stream1", data=data_array)
    group.attrs["timestamp"] = timestamp

# 读取数据
with h5py.File("data.h5", "r") as f:
    for group_name in f:
        group = f[group_name]
        loaded_image = group["image"][()]
        loaded_stream1 = group["stream1"][()]
        loaded_timestamp = group.attrs["timestamp"]
  • 优势:支持无损压缩,可高效管理数千条以上的记录,适合专业数据存储场景。

3. 图像单独存储+元数据轻量记录

适合需要直接用图像浏览器查看图像的场景,把图像存为JPG/PNG,元数据用CSV或SQLite记录。

  • 实现代码:
import cv2
import csv
import time
import os
import numpy as np

cam = cv2.VideoCapture(0)
result, image = cam.read()
data_array = np.zeros((2,3,4))
timestamp = time.time()

# 保存图像为JPG(体积远小于JSON存储)
img_filename = f"image_{int(timestamp)}.jpg"
cv2.imwrite(img_filename, image)

# 元数据写入CSV
data_row = [
    timestamp, 
    img_filename, 
    ",".join(map(str, data_array.flatten()))
]
file_exists = os.path.exists("metadata.csv")
with open("metadata.csv", "a", newline="") as f:
    writer = csv.writer(f)
    if not file_exists:
        writer.writerow(["timestamp", "image_filename", "stream1_data"])
    writer.writerow(data_row)

# 读取还原
with open("metadata.csv", "r") as f:
    reader = csv.DictReader(f)
    for row in reader:
        ts = float(row["timestamp"])
        img = cv2.imread(row["image_filename"])
        stream1 = np.array(list(map(float, row["stream1_data"].split(",")))).reshape(2,3,4)
  • 优势:图像可直接浏览,元数据文件体积极小,读写速度极快。

内容的提问来源于stack exchange,提问作者Chorum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:50:42