You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将编码字符串转为可读格式?Python3处理mat文件字符串数组方法

使用scipy.io.loadmat导入mat文件时,numpy.ndarray类型的字符串数组转可读格式(Python3)

用scipy.io.loadmat加载MATLAB的.mat文件时,里面的字符串数组(MATLAB里的cell数组)经常会被转换成嵌套的numpy.ndarray,元素多是bytes类型,得批量处理才能变成可读的字符串数组。我分两种常见情况给你解决方案:

情况1:一维/多维的直接bytes数组

如果你的数组里每个元素都是直接的bytes(比如形状是(n,)或者(n,m)的ndarray,元素是b'xxx'),可以用列表推导+numpy.reshape快速转换:

import scipy.io
import numpy as np

# 加载mat文件
mat_data = scipy.io.loadmat('your_file.mat')
# 取出字符串数组,假设变量名为str_array
str_array = mat_data['str_array']

# 先拉平数组解码,再还原原形状
decoded_array = np.array([x.decode('utf-8') for x in str_array.flatten()]).reshape(str_array.shape)

情况2:嵌套的单元素数组

有时候MATLAB的cell数组转过来后,每个元素是一个单元素的ndarray(比如array([b'xxx'], dtype='|S10')),这时候需要先取出里面的bytes再解码:

# 比如str_array的形状是(n, m),每个元素是单元素数组
decoded_array = np.empty(str_array.shape, dtype=str)
# 遍历每个位置解码
for i in range(str_array.shape[0]):
    for j in range(str_array.shape[1]):
        decoded_array[i, j] = str_array[i, j][0].decode('utf-8')

嫌循环麻烦的话,也可以用numpy.vectorize做向量化处理(注意:大数据量下性能不如循环,小数据量完全没问题):

decode_func = lambda x: x[0].decode('utf-8')
decoded_array = np.vectorize(decode_func)(str_array)

同样,要是你的字符串用的不是UTF-8编码,把decode('utf-8')换成对应的编码就行,比如decode('gbk')。


内容的提问来源于stack exchange,提问作者Vladislav Gladkikh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:02:10