将含数组的动态MF4信号保存为嵌套结构.mat文件
问题背景
需要从MF4文件提取信号存储为.mat格式的嵌套结构体,信号名以点分隔命名,示例如下:
camera_detection.timestamp.seconds.value camera_detection.timestamp.nanoseconds.value camera_detection.object._0_.type camera_detection.object._0_.position_x camera_detection.object._0_.position_y camera_detection.object._1_.type camera_detection.object._1_.position_x camera_detection.object._1_.position_y
目标.mat文件内的结构体格式要求:
struct camera_detection with fields: timestamp [1x1 struct with fields seconds and nanoseconds] object [2x1 struct of array with fields type, position_x, position_y]
当前使用asammdf读取MF4文件,scipy.io.savemat写入.mat文件,现有代码可正确处理普通嵌套字段(如时间戳信号),但无法自动识别_<数字>_格式的数组索引,无法生成结构体数组。要求结构完全动态构建,不预先获知信号结构、数组大小,支持任意层级的数组嵌套。
现有代码如下:
### output dictionary output = {} ### Read file mdf = MDF("file.mf4") for sig in mdf.iter_channels(): parts = sig.name.split(".") current = output.setdefault(parts[0], {}) for i in range(1,len(parts) - 1): current = current.setdefault(parts[i],{}) current[parts[-1]] = sig.samples ### Save to mat file savemat("out.mat", output)
实现方案
核心修改点:
- 通过正则识别路径中匹配
_<数字>_格式的段,解析为结构体数组的索引 - 遍历路径时动态调整节点类型:首次访问数组索引时,将对应节点从默认字典转换为列表,按索引填充空结构体
- 所有信号遍历完成后,递归将结构转换为
savemat兼容格式:Python字典对应MATLAB结构体,Python列表转换为dtype=object的N*1 numpy数组,对应MATLAB列方向结构体数组
完整代码:
import re import numpy as np from asammdf import MDF from scipy.io import savemat # 匹配_数字_格式的数组索引段 INDEX_PATTERN = re.compile(r'^_(\d+)_$') def convert_to_mat_compatible(obj): """递归将嵌套结构转换为scipy.io.savemat兼容格式""" if isinstance(obj, dict): # 字典对应MATLAB struct,递归处理所有字段 return {k: convert_to_mat_compatible(v) for k, v in obj.items()} elif isinstance(obj, list): # 列表对应MATLAB struct数组,转为n*1的object类型numpy数组 converted_items = [convert_to_mat_compatible(item) for item in obj] mat_arr = np.empty((len(converted_items), 1), dtype=object) for idx, item in enumerate(converted_items): mat_arr[idx, 0] = item return mat_arr else: # 信号采样值等基础类型直接返回 return obj # 初始化输出字典 output = {} mdf = MDF("file.mf4") for sig in mdf.iter_channels(): parts = sig.name.split(".") current = output parent = None current_key = None # 遍历除最后一个字段外的所有路径段 for part in parts[:-1]: index_match = INDEX_PATTERN.match(part) if index_match: # 当前段是数组索引,current必须为列表类型 idx = int(index_match.group(1)) # 首次访问时current是默认创建的字典,转换为列表 if isinstance(current, dict): current = [] parent[current_key] = current # 补全列表长度到索引+1,缺失位置填充空字典 while len(current) <= idx: current.append({}) # 进入索引对应位置的结构体 parent = current current_key = idx current = current[idx] else: # 当前段是普通字段名,current必须为字典类型 if part not in current: current[part] = {} parent = current current_key = part current = current[part] # 叶子节点赋值,与原逻辑一致 current[parts[-1]] = sig.samples # 转换为mat兼容格式后保存 mat_output = convert_to_mat_compatible(output) savemat("out.mat", mat_output, long_field_names=True)
说明
long_field_names=True参数用于支持超过31字符的字段名,兼容MATLAB 7.6及以上版本- 自动识别任意层级的数组嵌套,例如
a._0_.b._2_.c._1_.value这类多层数组的信号名可正确解析 - 数组长度根据遍历到的最大索引自动扩展,无需提前预知数组大小
- 普通嵌套字段的处理逻辑与原有代码完全一致,不会影响已可正常保存的信号
内容的提问来源于stack exchange,提问作者Anonymous
相关产品推荐
相关产品推荐

