如何用Python读取Fortran生成的无格式二进制文件?
Fortran无格式二进制文件的Python正确读取方法
问题背景
用户通过以下Fortran代码生成无格式二进制文件myFile.xyz:
PROGRAM main_program implicit none open (10, form='unformatted', file='myFile.xyz') write(10) 1 write(10) 400, 400, 1 write(10) 1.0d0, 1.25d0, 1.50d0, 1.75d0, 2.0d0, 2.25d0 write(10) 3.0d0, 3.25d0, 3.50d0, 3.75d0, 4.0d0, 4.25d0, 4.50d0, 4.75d0, 5.00d0 close(10) END
尝试用Python的numpy.fromfile直接读取时,因不了解Fortran无格式文件的特性导致读取失败。
核心原因
Fortran的unformatted无格式二进制文件会为每个write语句生成一个记录,每个记录的首尾会添加一个整数(通常为4字节int32,部分64位环境下为8字节int64),标记该记录的字节长度。用户的代码中有4个write,对应4个独立记录,每个记录都包含首尾的长度标记。
原Python代码错误地将所有数据视为float64类型读取,而实际上记录开头的长度标记是整数类型,后续数据的类型也与写入时一致(前两个记录是整数,后两个是双精度浮点数),导致类型不匹配,读取结果混乱。
正确读取代码
方法一:按记录逐个读取(推荐)
import numpy as np file_path = "myFile.xyz" with open(file_path, "rb") as f: # 读取第一个记录:单个整数1 rec_len = np.fromfile(f, dtype=np.int32, count=1)[0] int1 = np.fromfile(f, dtype=np.int32, count=1)[0] _ = np.fromfile(f, dtype=np.int32, count=1) # 读取第二个记录:三个整数400, 400, 1 rec_len = np.fromfile(f, dtype=np.int32, count=1)[0] ints = np.fromfile(f, dtype=np.int32, count=3) _ = np.fromfile(f, dtype=np.int32, count=1) # 读取第三个记录:6个双精度浮点数 rec_len = np.fromfile(f, dtype=np.int32, count=1)[0] floats1 = np.fromfile(f, dtype=np.float64, count=6) _ = np.fromfile(f, dtype=np.int32, count=1) # 读取第四个记录:9个双精度浮点数 rec_len = np.fromfile(f, dtype=np.int32, count=1)[0] floats2 = np.fromfile(f, dtype=np.float64, count=9) _ = np.fromfile(f, dtype=np.int32, count=1) # 合并结果 integers_data = np.concatenate([[int1], ints]) float_values = np.concatenate([floats1, floats2]) print("提取的整数数据:", integers_data) print("提取的双精度浮点数数据:", float_values)
注意事项
- 如果运行时出现读取长度不匹配的错误,说明你的环境中记录长度用的是8字节整数,将代码中的
np.int32替换为np.int64即可。 - 每个记录的长度标记值等于该记录内容的字节数,比如第一个记录内容是1个
int32(4字节),所以记录长度标记为4。
输出结果
运行上述代码后,会得到正确的提取结果:
提取的整数数据: [ 1 400 400 1] 提取的双精度浮点数数据: [1. 1.25 1.5 1.75 2. 2.25 3. 3.25 3.5 3.75 4. 4.25 4.5 4.75 5. ]
内容的提问来源于stack exchange,提问作者Prof. Dumbledore
相关产品推荐
相关产品推荐

