Python优雅实现Wolfram Mathematica输出转9列基础数组格式
Mathematica导出数组文件转9列纯文本的纯Python实现
需求概述
- 输入:Wolfram Mathematica导出的
.m格式数据文件,内容为嵌套大括号包裹的浮点数组,文件头包含Mathematica自动生成的注释行,数值可能使用Mathematica专属的*^标记科学计数法 - 输出:每行固定9个空格分隔浮点数的纯文本文件,格式与原有shell处理方案输出完全一致
- 约束:不依赖shell命令、不生成临时中间文件,纯Python实现,逻辑简洁可维护
输入文件内容示例:
(* Created with the Wolfram Language for Students - Personal Use Only : www.wolfram.com *) {{0.29344728841663786, 0.00037262711145454893, 0.7061800844719075, 67.41431300170986, 1.3887122472912174, 0.0014182932914303275, 500.97644711373647, 0.0002565333937360516, 105.86185844804378}, {0.29479428399557506, 0.0007813301223490133, 0.7044243858820759, 67.40475060370453, 1.3779372193629575, 0.00006103376259459755, 500.30876628350757, 0.00001106337484454747, 101.39952463245301}, {...
原有实现的问题
原有方案通过拼接多段shell命令做字符替换,需要生成临时文件,跨平台兼容性差(例如macOS默认sed与GNU sed参数不兼容,需额外安装gsed),逻辑零散难以维护:
# Convert chain.m to final_array.txt os.system("cat chain.m | tr '},' '\n' | tr '{{' ' ' | tr '{' ' ' | tr '}}' ' ' | gsed 's/\*\^-/e-/g' | gsed 's/\*\^/e/g' | grep -v '(' > out.txt") a=np.loadtxt('out.txt') os.system('rm -f out.txt') nline = int(len(a)/9) b=np.reshape(a,(nline,9)) np.savetxt('final_array.txt', b)
目标输出格式示例:
0.29344728841663786 0.00037262711145454893 0.7061800844719075 67.41431300170986 1.3887122472912174 0.0014182932914303275 500.97644711373647 0.0002565333937360516 105.86185844804378 0.29479428399557506 0.0007813301223490133 0.7044243858820759 67.40475060370453 1.3779372193629575 0.00006103376259459755 500.30876628350757 0.00001106337484454747 101.39952463245301
实现方案
方案1:仅依赖Python标准库(无第三方依赖)
核心逻辑:读取文件后先移除Mathematica注释、转换科学计数法格式,再通过正则直接提取所有浮点数值,按9个一组重组写入文件,鲁棒性强,不会因为换行、多余空格、括号位置变化导致解析失败。
import re # 配置参数 INPUT_PATH = "chain.m" OUTPUT_PATH = "final_array.txt" COL_COUNT = 9 # 读取文件内容 with open(INPUT_PATH, "r", encoding="utf-8") as f: raw_content = f.read() # 清洗内容 # 1. 移除所有Mathematica块注释 cleaned = re.sub(r"\(\*.*?\*\)", "", raw_content, flags=re.DOTALL) # 2. 将Mathematica科学计数法标记*^转换为通用e标记 cleaned = cleaned.replace("*^", "e") # 3. 提取所有浮点数 all_nums = list(map(float, re.findall(r"-?\d+\.?\d*(?:e[+-]?\d+)?", cleaned))) # 按固定列数写入输出 with open(OUTPUT_PATH, "w", encoding="utf-8") as f: for row_idx in range(0, len(all_nums), COL_COUNT): row_values = all_nums[row_idx:row_idx+COL_COUNT] f.write(" ".join(map(str, row_values)) + "\n")
方案2:适配numpy工作流(与原有输出100%兼容)
如果后续需要用numpy处理数据,可以直接用numpy的IO接口完成解析和保存,代码更简洁:
import re import numpy as np INPUT_PATH = "chain.m" OUTPUT_PATH = "final_array.txt" COL_COUNT = 9 with open(INPUT_PATH, "r", encoding="utf-8") as f: raw_content = f.read() cleaned = re.sub(r"\(\*.*?\*\)", "", raw_content, flags=re.DOTALL).replace("*^", "e") # 将所有括号、逗号替换为空格,直接用numpy解析数值数组 all_nums = np.fromstring(re.sub(r"[{},]", " ", cleaned), sep=" ") # 自动按列数重组,不需要手动计算行数 data_array = all_nums.reshape(-1, COL_COUNT) # 保存为空格分隔文本,与原有np.savetxt输出格式完全一致 np.savetxt(OUTPUT_PATH, data_array)
方案优势
- 全平台兼容,不需要依赖sed、tr等shell命令,也不需要安装额外系统工具
- 无临时文件生成,减少不必要的IO操作
- 解析逻辑容错性高,自动忽略多余空格、换行、不同位置的括号分隔符
- 自动识别并转换Mathematica格式的科学计数法数值,自动过滤所有注释内容
内容的提问来源于stack exchange,提问作者user1773603
相关产品推荐
相关产品推荐

