如何在Python中读取MATLAB保存的字典数据?
问题描述
我用MATLAB创建了字典,将字典的keys和values分别用-v7.3参数保存为.mat文件,MATLAB代码如下:
cnn6 = load("imgnet19k_less_feature_map_20230412_alexnet_cnn6_dict.mat"); dictionary_analyze = cnn6.dictionary_imagenames_features; keys_in_order = keys(dictionary_analyze); values_in_order = values(dictionary_analyze); save("cnn6_keys_in_order_imagenet_19k_v2.mat", "keys_in_order", '-v7.3'); save("cnn6_values_in_order_imagenet_19k_v2.mat", "values_in_order", '-v7.3');
字典结构说明:
- 字典的key为字符串类型(对应图片名称)
- 每个value是二维数值数组(特征向量)
我尝试用Python读取这些.mat文件,代码如下:
import os import sys import pickle import numpy as np import h5py import tables from scipy.io import loadmat normal_semantic_features_dir = "./normal_semantic_features/" alexnet_features = ['cnn2', 'cnn4', 'cnn6', 'cnn8'] #dict_cnns = h5py.File(normal_semantic_features_dir + 'cnn2_keys_in_order_imagenet_19k_v2.mat', "r") f1 = h5py.File(normal_semantic_features_dir + 'cnn8_keys_in_order_imagenet_19k_v2.mat', "r") f2 = h5py.File(normal_semantic_features_dir + 'cnn8_values_in_order_imagenet_19k_v2.mat', "r") array_of_keys = dict_cnns['keys_in_order'] print(list(f1.keys()))
运行结果:
> ['#refs#', '#subsystem#', 'keys_in_order'] a = f1["keys_in_order"] type(a) > h5py._hl.dataset.Dataset for k, v in annots.items(): print(k,"........" ,annots[k]) print("....") > __header__ ........ b'MATLAB 5.0 MAT-file, Platform: PCWIN64, Created on: Fri Apr 14 13:32:47 2023' .... > __version__ ........ 1.0 .... > __globals__ ........ [] .... None ........ [(b'keys_in_order', b'MCOS', b'string', array([[3707764736], > [ 2], > [ 1], > [ 1], > [ 1], > [ 1]], dtype=uint32))] .... > __function_workspace__ ........ [[ 0 1 73 ... 0 0 0]]
请问如何读取这些.mat文件内容并转换为Python风格的字典?
解决方案
MATLAB的-v7.3格式本质是HDF5格式,用h5py读取时,MATLAB的字符串数组、cell数组会以引用形式存储,需要特殊解析。以下是完整的读取和转换代码:
完整实现代码
import h5py import numpy as np def read_matlab_strings(h5_file, dataset_name): """解析MATLAB v7.3文件中的字符串数组""" ref_dataset = h5_file[dataset_name] # 获取所有字符串引用并扁平化 str_refs = ref_dataset[()].flatten() keys = [] for ref in str_refs: # 读取引用指向的字节数据,解码为UTF-8字符串 str_bytes = h5_file[ref][()].tobytes() keys.append(str_bytes.decode('utf-8')) return keys def read_matlab_values(h5_file, dataset_name): """解析MATLAB v7.3文件中的数值数组集合""" val_dataset = h5_file[dataset_name] val_refs = val_dataset[()].flatten() values = [] for ref in val_refs: # 读取引用指向的数组,转置适配Python行优先存储 arr = np.array(h5_file[ref[0]]).T values.append(arr) return values # 文件路径配置 normal_semantic_features_dir = "./normal_semantic_features/" key_file = f"{normal_semantic_features_dir}cnn8_keys_in_order_imagenet_19k_v2.mat" val_file = f"{normal_semantic_features_dir}cnn8_values_in_order_imagenet_19k_v2.mat" # 读取并转换为Python字典 with h5py.File(key_file, 'r') as f_keys, h5py.File(val_file, 'r') as f_vals: keys_list = read_matlab_strings(f_keys, 'keys_in_order') values_list = read_matlab_values(f_vals, 'values_in_order') python_dict = dict(zip(keys_list, values_list)) # 验证结果示例 print("前3个key:", keys_list[:3]) print("第一个value的形状:", python_dict[keys_list[0]].shape)
代码说明
- 读取字符串keys:MATLAB的字符串在v7.3文件中以引用形式存储,需要遍历每个引用,读取对应的字节数据并解码为UTF-8字符串。
- 读取数值values:字典的values是嵌套的数值数组,同样以引用形式存储,读取后转置数组(MATLAB是列优先存储,Python为行优先)。
- 构建Python字典:用
zip()将keys和values一一配对,直接转换为原生Python字典。
注意事项
- 确保依赖库已安装:执行
pip install h5py numpy - 如果values是其他类型的嵌套结构,需要根据MATLAB原数据结构调整
read_matlab_values的解析逻辑。
内容的提问来源于stack exchange,提问作者Kadaj13
相关产品推荐
相关产品推荐

