如何解决使用np.save保存数据时的非均匀形状错误?
问题:使用np.save保存数据时触发形状不兼容错误
尝试用np.save将数据保存为.npy文件时,出现以下错误:
setting an array element with a sequence. The requested array has an inhomogeneous shape after 2 dimensions. The detected shape was (200, 2) + inhomogeneous part.
相关代码
dataset = [] # 循环内逻辑 img = cv2.imread(image_file,1) img = cv2.resize(img,(224,224)) img = img.astype(np.float32) locations = [value1, value2, value3, value4] # 1x4整数向量 dataset.append(np.array(img), np.array(locations), predicted_class, filename) # predicted_class为整数,filename为字符串
代码逻辑:遍历图片集,处理图片后生成1x4向量,将图片数组、locations数组、预测类别、文件名加入dataset列表。
检查dataset结构,每个元素格式为:
[ array([[[#,#,#],...,[#,#,#]]], dtype=float32), array([#,#,#,#]), #, 'filename.jpg']
执行保存代码时触发错误:
np.save('dataset.npy',dataset)
错误回溯:
ValueError Traceback (most recent call last) <ipython-input-3-1b033c5c2e47> in <cell line: 4>() 2 print(dataset[0]) 3 os.chdir('/content/drive/My Drive/Misc/') ----> 4 np.save('dataset.npy', dataset) /usr/local/lib/python3.10/dist-packages/numpy/lib/npyio.py in save(file, arr, allow_pickle, fix_imports) 543 544 with file_ctx as fid: ---> 545 arr = np.asanyarray(arr) 546 format.write_array(fid, arr, allow_pickle=allow_pickle, 547 pickle_kwargs=dict(fix_imports=fix_imports)) ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 2 dimensions. The detected shape was (200, 2) + inhomogeneous part.
解决方案
错误原因
np.save要求输入为同构numpy数组,但你的dataset是混合了不同类型、不同形状元素的列表:每个元素包含(224,224,3)的图片数组、(4,)的locations数组、整数、字符串。当np.save尝试将该列表转为numpy数组时,无法构造统一形状的数组,因此触发形状不兼容错误。
可行解决方法
方案1:用pickle保存异构列表
因为数据是混合类型,Python原生的pickle更适合保存这种异构结构:
import pickle # 保存 with open('dataset.pkl', 'wb') as f: pickle.dump(dataset, f) # 读取 with open('dataset.pkl', 'rb') as f: loaded_dataset = pickle.load(f)
方案2:拆分数据为同构数组分别保存
将不同类型/形状的数据拆分,各自保存为独立的.npy文件:
# 拆分数据 images = np.array([item[0] for item in dataset]) locations = np.array([item[1] for item in dataset]) predicted_classes = np.array([item[2] for item in dataset]) filenames = np.array([item[3] for item in dataset], dtype=np.str_) # 分别保存 np.save('images.npy', images) np.save('locations.npy', locations) np.save('predicted_classes.npy', predicted_classes) np.save('filenames.npy', filenames) # 读取示例 loaded_images = np.load('images.npy') loaded_locations = np.load('locations.npy')
额外修正:代码中的语法错误
原代码中dataset.append的写法有误,append仅接受一个参数,需将四个元素打包为元组或列表再添加:
# 修正后的append代码 dataset.append( (np.array(img), np.array(locations), predicted_class, filename) )
内容的提问来源于stack exchange,提问作者August Phang
相关产品推荐
相关产品推荐

