Python脚本耦合CMG油藏模拟:大输出文件读取优化问询
问题描述
我正尝试将Python脚本与油气行业的CMG油藏模拟软件耦合,该软件每个时间步(如0.1天、0.2天等)生成输出文件。我编写了一个函数,用于读取输出文件、提取CO₂摩尔分数、计算各时间步的增量值,后续将这些值用于与软件耦合,使其根据各时间步的CO₂摩尔分数执行特定命令。
目前遇到的问题是:该函数每个时间步都会从头读取整个输出文件,随着文件不断增大(可达数百万行),操作耗时极高。我希望优化代码,让下一个时间步能从上次提取数据的位置开始搜索,而非从头开始。
现有函数代码
def co2fromOUT(tStart): #tStart = 2 import os t_find = "TIME: " + str(tStart); #+ " days"; # from column 1 to 12 in CMG input data file str2find_1 = "Compositions (mole fractions)"; # from column 47 to 84 in CMG input data file str2find_2 = "CO2" #################################### 2 locations - This file & python_GEM one path_to_file = 'C:/Users/name/Desktop/test 1/run2/2.OUT' file_name = "2" + "_" + "molCO2" + ".data" # file to save moles of CO2 init = str(0) + "\t" + str(0) + "\t" + str(0) with open(path_to_file, 'r') as f: lines = f.readlines() # Create a new file to write CO2 moles co2store = open(file_name,"a+") if os.stat(file_name).st_size == 0: co2store.write(init) co2store.close() count = 0 curr_molCO2 = 0 incr_molCO2 = 0 for i in range(0, len(lines)): res1 = lines[i].find(t_find) # search for t_find in line i if res1 == -1: # if t_find is not found in current line continue else: # Case for solubility res2 = lines[i - 12].find(str2find_1) # find str2find_1 in line i-12 if res2 == -1: print("NOT FOUND 11 lines behind current time - Compositions (mole fractions)") continue else: # Case for solubility res3 = lines[i - 7].find(str2find_2) # find str2find_2 in line i-7 if res3 == -1: print("NOT FOUND 6 lines behind current time - CO2") continue else: with open(file_name,'r') as co2stores: for co2line in co2stores: pass #last_line = lin #co2line = co2stores.readlines() #print(co2line) prev_molCO2 = co2line.split("\t") #print(prev_molCO2) #print(prev_molCO2[1]) #lst = line[i - 24].strip().split() #print( lines[i - 24][42:54] ) curr_molCO2 = float( lines[i - 7][42:55] ) # CO2 moles appear 6 lines below str2find_1 incr_molCO2 = curr_molCO2 - float( prev_molCO2[1] ) molCO2store = str(tStart) + "\t" + str(curr_molCO2) + "\t" + str(incr_molCO2) co2store = open(file_name,"a+") co2store.write("\n" + molCO2store) co2store.close() return incr_molCO2, curr_molCO2
部分输出文件示例
Stream Cum Inj Cum Prod ----- ------- -------- Oil 0.00000E+00 0.00000E+00 bbl Gas 9.37050E+07 6.24700E+07 cuft Wet Gas 0.00000E+00 6.26270E+07 cuft Water 9.29089E+01 2.22459E+02 bbl Inj Rate Prod Rate Phase moles/d moles/d ----- ----------- ----------- Oil 0.00000E+00 0.00000E+00 Gas 1.80444E+05 1.19860E+05 Wet Gas 0.00000E+00 1.19860E+05 Water 0.00000E+00 8.61632E+03 Compositions (mole fractions) of Total Field Produced and Injected Streams: Oil Gas Wet Gas Component Produced Injected Produced Injected Produced Injected --------- -------- -------- -------- -------- -------- -------- CO2 0.00000E+00 0.00000E+00 1.43080E-05 1.00000E+00 1.43080E-05 0.00000E+00 C1 0.00000E+00 0.00000E+00 9.99986E-01 0.00000E+00 9.99986E-01 0.00000E+00 C2 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 C3 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 CO2_T 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 0.00000E+00 1 *********************************************************************************************************************************** TIME: 624.7 days G E M S E C T O R S U M M A R Y DATE: 1981:09:16
优化方案
核心思路是记录上次读取的字节位置,避免每次从头加载全文件,同时用状态机模式解析输出,替代不可靠的行号回溯。
1. 优化后的代码
import os def co2fromOUT(tStart): t_find = f"TIME: {tStart}" str2find_1 = "Compositions (mole fractions)" str2find_2 = "CO2" path_to_file = 'C:/Users/name/Desktop/test 1/run2/2.OUT' file_name = "2_molCO2.data" pos_file = "read_position.txt" # 保存上次读取的字节位置 # 初始化CO2存储文件 if not os.path.exists(file_name) or os.stat(file_name).st_size == 0: with open(file_name, "w") as f: f.write("0\t0\t0") # 获取上次读取的起始位置 start_pos = 0 if os.path.exists(pos_file): with open(pos_file, "r") as f: content = f.read().strip() if content.isdigit(): start_pos = int(content) curr_molCO2 = 0 incr_molCO2 = 0 in_time_block = False # 是否进入目标时间步的块 comp_block_found = False # 是否找到成分数据块 co2_line_counter = 0 # 成分块内的行计数器 # 从上次位置开始逐行读取文件 with open(path_to_file, 'r') as f: f.seek(start_pos) for line in f: current_pos = f.tell() # 记录当前读取的字节位置 # 匹配目标时间步 if t_find in line: in_time_block = True comp_block_found = False co2_line_counter = 0 # 在目标时间步内查找成分数据块 elif in_time_block: if str2find_1 in line: comp_block_found = True co2_line_counter = 0 # 在成分数据块内找CO2行 elif comp_block_found: co2_line_counter += 1 # 根据输出格式,CO2行在成分标题后第5行(含空行) if co2_line_counter == 5 and str2find_2 in line: # 分割行内容提取CO2摩尔分数(比固定列索引更可靠) line_parts = line.strip().split() if len(line_parts) >= 2: curr_molCO2 = float(line_parts[1]) # 取Produced列的CO2值,可根据需求调整索引 # 获取上一次的CO2值 with open(file_name, 'r') as co2_f: last_line = co2_f.readlines()[-1].strip() prev_molCO2 = float(last_line.split("\t")[1]) incr_molCO2 = curr_molCO2 - prev_molCO2 # 写入新数据 with open(file_name, "a") as co2_f: co2_f.write(f"\n{tStart}\t{curr_molCO2}\t{incr_molCO2}") # 更新读取位置,避免重复处理 with open(pos_file, "w") as f_pos: f_pos.write(str(current_pos)) # 找到目标数据后退出循环 break # 实时更新读取位置,防止中途退出丢失进度 with open(pos_file, "w") as f_pos: f_pos.write(str(current_pos)) return incr_molCO2, curr_molCO2
2. 关键优化点
- 字节偏移记录:用
f.tell()获取当前读取的字节位置,下次通过f.seek(start_pos)直接跳转,彻底避免从头读取全文件。 - 状态机解析:跟踪当前处于「时间步块」「成分数据块」的状态,按顺序查找目标数据,替代原代码中不可靠的行号回溯(
i-12/i-7),适配格式微小变化。 - 逐行读取:不再一次性加载数百万行到内存,内存占用极低,适合超大文件。
- 减少IO操作:合并文件打开/关闭逻辑,读取存储文件时直接取最后一行,避免循环遍历所有行。
3. 注意事项
- 确保CMG输出文件是追加写入模式,即每个时间步在文件末尾添加内容,否则字节偏移记录会失效。
- 若输出文件被覆盖,需手动删除
read_position.txt重置读取位置。 - 提取CO2值时,用
split()分割行内容比固定列索引更可靠,可根据实际输出格式调整line_parts的索引。
内容的提问来源于stack exchange,提问作者user621895
相关产品推荐
相关产品推荐

