使用scipy.io.loadmat加载4GB SWAN输出.mat文件时遇读取错误
加载SWAN输出的4GB .mat文件时出现ValueError: Did not read any bytes
我尝试用scipy的scipy.io.loadmat加载一个4GB的Matlab文件,该文件是波浪数值模型SWAN的输出,文件大小显示正常,但加载时报错ValueError: Did not read any bytes。以下是我使用mhkit库的相关代码,求解决思路。
def _read_block_mat(swan_file): ''' Reads in SWAN matlab output and creates a dictionary of DataFrames for each swan output variable. Parameters ---------- swan_file: str filename to import Returns ------- dataDict: Dictionary Dictionary of DataFrame of swan output variables ''' assert isinstance(swan_file, str), 'swan_file must be of type str' assert isfile(swan_file)==True, f'File not found: {swan_file}' dataDict = loadmat(swan_file, struct_as_record=False, squeeze_me=True) removeKeys = ['__header__', '__version__', '__globals__'] for key in removeKeys: dataDict.pop(key, None) for key in dataDict.keys(): dataDict[key] = pd.DataFrame(dataDict[key]) return dataDict def read_block(swan_file): ''' Reads in SWAN block output with headers and creates a dictionary of DataFrames for each SWAN output variable in the output file. Parameters ---------- swan_file: str swan block file to import Returns ------- data: Dictionary Dictionary of DataFrame of swan output variables metaDict: Dictionary Dictionary of metaData dependent on file type ''' assert isinstance(swan_file, str), 'swan_file must be of type str' assert isfile(swan_file)==True, f'File not found: {swan_file}' extension = swan_file.split('.')[1].lower() if extension == 'mat': dataDict = _read_block_mat(swan_file) metaData = {'filetype': 'mat', 'variables': [var for var in dataDict.keys()]} else: dataDict, metaData = _read_block_txt(swan_file) return dataDict, metaData
可能的解决思路
- 验证文件完整性:大文件传输或存储时容易损坏,计算文件的哈希值(如MD5、SHA256),和原始文件的哈希值对比,确认文件未损坏。
- 针对大文件适配读取方式:
scipy.io.loadmat默认一次性加载整个文件到内存,4GB文件可能超出内存限制或触发读取异常。可以尝试:- 使用
mat73库,它支持mat7.3格式的大文件分块读取,无需一次性加载全部数据。 - 用Matlab打开原文件,将其拆分成多个小.mat文件后再逐个读取。
- 使用
- 确认Mat文件版本兼容性:SWAN输出的.mat可能是高版本(如mat7.3),而
scipy.io.loadmat对该版本支持有限。可以:- 用Matlab打开文件,另存为低版本(如v7)后再尝试读取。
- 直接用
h5py读取(因为mat7.3本质是HDF5格式),示例代码:import h5py with h5py.File('your_swan_file.mat', 'r') as f: # 查看文件内的所有数据集 print("Available variables:", list(f.keys())) # 提取指定变量数据 example_data = f['your_variable_name'][()]
- 检查文件路径:确保文件路径没有特殊字符、空格或过长的路径,避免读取时路径解析出错。
- 排查内存资源:加载4GB文件需要足够的内存空间,关闭其他占用大量内存的程序,释放内存后重试。
内容的提问来源于stack exchange,提问作者Edouardo
相关产品推荐
相关产品推荐

