如何用Fortran从输入文件指定行开始读取数据?
直接从超大Fortran输入文件的指定行读取数据
问题背景
处理超大文本输入文件时,需要从中间某行(如95038000行)开始读取,但当前通过逐行跳过的方式(如下代码)效率极低,耗时过长:
implicit none integer :: i open(99, file = 'input.dat') do i=1, 10000 read(99,*) end do ! skips the first 10,000 lines of file input.dat
输入文件为结构化文本,包含重复的块结构,示例如下:
100 1 0.01 20000 20000 He 51.71286 -72.51866 -18.82236 He 26.74500 -55.83966 -21.50548 ... 100 2 0.01 20000 20000 He -75.32897 -18.60672 25.41119 ...
解决方案
Fortran本身没有直接按行号跳转的内置函数,但可以根据文件特点选择以下高效方法:
1. 固定行宽文件:直接访问模式定位
如果文件每行的字节数完全固定,可通过直接访问模式打开文件,直接定位到目标行:
implicit none integer, parameter :: SKIP_LINES = 95038000 integer :: unit, rec_len, status character(len=256) :: dummy_line ! 先获取每行的记录长度 open(newunit=unit, file='input.dat', status='old', action='read') inquire(unit=unit, recl=rec_len) close(unit) ! 以直接访问模式打开,直接跳转到目标行的下一行 open(newunit=unit, file='input.dat', status='old', action='read', access='direct', recl=rec_len) read(unit, rec=SKIP_LINES+1, iostat=status) dummy_line if (status /= 0) error stop "Failed to locate target line" ! 后续正常读取数据 ! read(unit, rec=...) your_variables close(unit)
注意:此方法仅适用于所有行字节数完全一致的文件,否则定位会出错。
2. 结构化文件:利用块标记定位
针对你提供的重复块结构文件,可直接搜索块标记(如100或块编号1/2),无需计算具体行数:
implicit none integer :: unit, block_id, target_block = 2, status character(len=256) :: line open(newunit=unit, file='input.dat', status='old', action='read') ! 循环搜索目标块 do read(unit, '(A)', iostat=status) line if (status /= 0) error stop "File ended before finding target block" ! 匹配块开头的"100"标记 if (line(1:3) == '100') then ! 读取下一行的块编号 read(unit, *, iostat=status) block_id if (block_id == target_block) exit ! 找到目标块,停止搜索 end if end do ! 开始读取目标块内的He数据 do read(unit, *, iostat=status) line ! 替换为实际变量读取(如elem, x, y, z) if (status /= 0) exit print *, line ! 示例输出,替换为你的数据处理逻辑 end do close(unit)
此方法效率远高于逐行跳过,因为它直接跳过无关块,无需解析每一行。
3. 系统工具预处理(Unix/Linux/Windows)
利用系统自带的文件处理工具提前裁剪文件,速度远快于Fortran逐行读取:
- Unix/Linux:执行
tail命令直接提取从目标行开始的内容:tail -n +95038001 input.dat > trimmed_input.dat - Windows:使用PowerShell或
more命令:
或Get-Content input.dat -Skip 95038000 > trimmed_input.datmore +95038000 input.dat > trimmed_input.dat
也可在Fortran中直接调用系统命令完成预处理:
implicit none character(len=256) :: cmd ! Unix/Linux版本 cmd = 'tail -n +95038001 input.dat > trimmed_input.dat' ! Windows PowerShell版本 ! cmd = 'powershell -Command "Get-Content input.dat -Skip 95038000 > trimmed_input.dat"' call system(cmd) ! 读取裁剪后的文件 open(99, file='trimmed_input.dat') ! 后续正常读取 close(99)
内容的提问来源于stack exchange,提问作者Adam.L
相关产品推荐
相关产品推荐

