You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python优雅实现Wolfram Mathematica输出转9列基础数组格式

Mathematica导出数组文件转9列纯文本的纯Python实现

需求概述

  • 输入:Wolfram Mathematica导出的.m格式数据文件,内容为嵌套大括号包裹的浮点数组,文件头包含Mathematica自动生成的注释行,数值可能使用Mathematica专属的*^标记科学计数法
  • 输出:每行固定9个空格分隔浮点数的纯文本文件,格式与原有shell处理方案输出完全一致
  • 约束:不依赖shell命令、不生成临时中间文件,纯Python实现,逻辑简洁可维护

输入文件内容示例:

(* Created with the Wolfram Language for Students - Personal Use Only : www.wolfram.com *)
{{0.29344728841663786, 0.00037262711145454893, 0.7061800844719075,
  67.41431300170986, 1.3887122472912174, 0.0014182932914303275,
  500.97644711373647, 0.0002565333937360516, 105.86185844804378},
 {0.29479428399557506, 0.0007813301223490133, 0.7044243858820759,
  67.40475060370453, 1.3779372193629575, 0.00006103376259459755,
  500.30876628350757, 0.00001106337484454747, 101.39952463245301},
{...

原有实现的问题

原有方案通过拼接多段shell命令做字符替换,需要生成临时文件,跨平台兼容性差(例如macOS默认sed与GNU sed参数不兼容,需额外安装gsed),逻辑零散难以维护:

# Convert chain.m to final_array.txt
os.system("cat chain.m | tr '},' '\n' | tr '{{' ' ' | tr '{' ' ' | tr '}}' ' ' | gsed 's/\*\^-/e-/g' | gsed 's/\*\^/e/g' | grep -v '(' > out.txt")
a=np.loadtxt('out.txt')
os.system('rm -f out.txt')
nline = int(len(a)/9)
b=np.reshape(a,(nline,9))
np.savetxt('final_array.txt', b)

目标输出格式示例:

0.29344728841663786 0.00037262711145454893 0.7061800844719075 67.41431300170986 1.3887122472912174 0.0014182932914303275 500.97644711373647 0.0002565333937360516 105.86185844804378
0.29479428399557506 0.0007813301223490133 0.7044243858820759 67.40475060370453 1.3779372193629575 0.00006103376259459755 500.30876628350757 0.00001106337484454747 101.39952463245301

实现方案

方案1:仅依赖Python标准库(无第三方依赖)

核心逻辑:读取文件后先移除Mathematica注释、转换科学计数法格式,再通过正则直接提取所有浮点数值,按9个一组重组写入文件,鲁棒性强,不会因为换行、多余空格、括号位置变化导致解析失败。

import re

# 配置参数
INPUT_PATH = "chain.m"
OUTPUT_PATH = "final_array.txt"
COL_COUNT = 9

# 读取文件内容
with open(INPUT_PATH, "r", encoding="utf-8") as f:
    raw_content = f.read()

# 清洗内容
# 1. 移除所有Mathematica块注释
cleaned = re.sub(r"\(\*.*?\*\)", "", raw_content, flags=re.DOTALL)
# 2. 将Mathematica科学计数法标记*^转换为通用e标记
cleaned = cleaned.replace("*^", "e")
# 3. 提取所有浮点数
all_nums = list(map(float, re.findall(r"-?\d+\.?\d*(?:e[+-]?\d+)?", cleaned)))

# 按固定列数写入输出
with open(OUTPUT_PATH, "w", encoding="utf-8") as f:
    for row_idx in range(0, len(all_nums), COL_COUNT):
        row_values = all_nums[row_idx:row_idx+COL_COUNT]
        f.write(" ".join(map(str, row_values)) + "\n")

方案2:适配numpy工作流(与原有输出100%兼容)

如果后续需要用numpy处理数据,可以直接用numpy的IO接口完成解析和保存,代码更简洁:

import re
import numpy as np

INPUT_PATH = "chain.m"
OUTPUT_PATH = "final_array.txt"
COL_COUNT = 9

with open(INPUT_PATH, "r", encoding="utf-8") as f:
    raw_content = f.read()

cleaned = re.sub(r"\(\*.*?\*\)", "", raw_content, flags=re.DOTALL).replace("*^", "e")
# 将所有括号、逗号替换为空格,直接用numpy解析数值数组
all_nums = np.fromstring(re.sub(r"[{},]", " ", cleaned), sep=" ")
# 自动按列数重组,不需要手动计算行数
data_array = all_nums.reshape(-1, COL_COUNT)
# 保存为空格分隔文本,与原有np.savetxt输出格式完全一致
np.savetxt(OUTPUT_PATH, data_array)

方案优势

  • 全平台兼容,不需要依赖sed、tr等shell命令,也不需要安装额外系统工具
  • 无临时文件生成,减少不必要的IO操作
  • 解析逻辑容错性高,自动忽略多余空格、换行、不同位置的括号分隔符
  • 自动识别并转换Mathematica格式的科学计数法数值,自动过滤所有注释内容

内容的提问来源于stack exchange,提问作者user1773603

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 16:48:17