Python读取ASCII文件速度逐渐变慢的问题排查与解决求助
处理仿真Verbose ASCII输出并转换为Pandas DataFrame的实践
最近我在处理一个仿真输出的ASCII格式结果文件,这个文件是用verbose模式生成的——每个时间步都会输出结构完全一致但数值更新的表格。为了后续能方便地绘制仿真变量的变化图表,我用Python把这些重复的表格转换成了pandas DataFrame,整个流程拆成了两个核心步骤:
第一步:快速拆分文件为时间步对应的片段
因为仿真文件可能很大,直接全量加载会占太多内存,所以我先做了快速遍历,定位每个时间步表格的起止标记,把大文件切割成和时间步数量相等的小片段:
- 核心逻辑是逐行扫描,识别每个时间步的起始和结束标识(比如我用的是
TIME STEP和END OF STEP,你得根据自己的文件改) - 遇到新的时间步就把上一个片段存起来,避免内存过载
对应的代码示例:
def split_simulation_file(file_path, start_marker="TIME STEP", end_marker="END OF STEP"): step_chunks = [] current_chunk = [] with open(file_path, 'r') as f: for line in f: stripped_line = line.strip() if start_marker in stripped_line: # 遇到新时间步,先保存上一个已收集的片段 if current_chunk: step_chunks.append(current_chunk) current_chunk = [] current_chunk.append(stripped_line) elif end_marker in stripped_line: current_chunk.append(stripped_line) step_chunks.append(current_chunk) current_chunk = [] else: if current_chunk: current_chunk.append(stripped_line) # 处理最后一个未被添加的片段 if current_chunk: step_chunks.append(current_chunk) return step_chunks
第二步:将每个时间步片段转换为DataFrame
拿到拆分后的片段后,接下来就是把每个片段里的表格提取出来转成DataFrame,还可以给每个DataFrame加上时间步索引,方便后续合并分析:
- 先过滤掉每个片段里的非表格内容(比如时间步说明、分隔线)
- 提取表头和数据行,转成DataFrame后统一添加时间步列,最后合并成一个大的DataFrame方便后续绘图
对应的代码示例:
import pandas as pd def chunks_to_dataframes(step_chunks): all_time_step_dfs = [] for step_idx, chunk in enumerate(step_chunks, start=1): # 过滤掉非表格行,这里要根据你的文件格式调整过滤规则 table_content = [line for line in chunk if not line.startswith(("TIME STEP", "END OF STEP"))] if not table_content: continue # 提取表头和数据行 header = table_content[0].split() data_rows = [line.split() for line in table_content[1:]] # 转成DataFrame并指定数据类型为float step_df = pd.DataFrame(data_rows, columns=header, dtype=float) step_df["time_step"] = step_idx all_time_step_dfs.append(step_df) # 合并所有时间步的DataFrame combined_df = pd.concat(all_time_step_dfs, ignore_index=True) return combined_df
小提示
实际使用时一定要根据自己的仿真文件格式调整标记和过滤规则!不同仿真工具的verbose输出差异很大——有的用连续的
=====做分隔,有的表头是固定的列名行,这些细节都得针对性修改才能准确提取表格内容。
内容的提问来源于stack exchange,提问作者Alemanio
相关产品推荐
相关产品推荐

