You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取ASCII文件速度逐渐变慢的问题排查与解决求助

处理仿真Verbose ASCII输出并转换为Pandas DataFrame的实践

最近我在处理一个仿真输出的ASCII格式结果文件,这个文件是用verbose模式生成的——每个时间步都会输出结构完全一致但数值更新的表格。为了后续能方便地绘制仿真变量的变化图表,我用Python把这些重复的表格转换成了pandas DataFrame,整个流程拆成了两个核心步骤:

第一步:快速拆分文件为时间步对应的片段

因为仿真文件可能很大,直接全量加载会占太多内存,所以我先做了快速遍历,定位每个时间步表格的起止标记,把大文件切割成和时间步数量相等的小片段:

  • 核心逻辑是逐行扫描,识别每个时间步的起始和结束标识(比如我用的是TIME STEP和END OF STEP,你得根据自己的文件改)
  • 遇到新的时间步就把上一个片段存起来,避免内存过载

对应的代码示例:

def split_simulation_file(file_path, start_marker="TIME STEP", end_marker="END OF STEP"):
    step_chunks = []
    current_chunk = []
    with open(file_path, 'r') as f:
        for line in f:
            stripped_line = line.strip()
            if start_marker in stripped_line:
                # 遇到新时间步,先保存上一个已收集的片段
                if current_chunk:
                    step_chunks.append(current_chunk)
                    current_chunk = []
                current_chunk.append(stripped_line)
            elif end_marker in stripped_line:
                current_chunk.append(stripped_line)
                step_chunks.append(current_chunk)
                current_chunk = []
            else:
                if current_chunk:
                    current_chunk.append(stripped_line)
    # 处理最后一个未被添加的片段
    if current_chunk:
        step_chunks.append(current_chunk)
    return step_chunks

第二步:将每个时间步片段转换为DataFrame

拿到拆分后的片段后,接下来就是把每个片段里的表格提取出来转成DataFrame,还可以给每个DataFrame加上时间步索引,方便后续合并分析:

  • 先过滤掉每个片段里的非表格内容(比如时间步说明、分隔线)
  • 提取表头和数据行,转成DataFrame后统一添加时间步列,最后合并成一个大的DataFrame方便后续绘图

对应的代码示例:

import pandas as pd

def chunks_to_dataframes(step_chunks):
    all_time_step_dfs = []
    for step_idx, chunk in enumerate(step_chunks, start=1):
        # 过滤掉非表格行,这里要根据你的文件格式调整过滤规则
        table_content = [line for line in chunk if not line.startswith(("TIME STEP", "END OF STEP"))]
        if not table_content:
            continue
        # 提取表头和数据行
        header = table_content[0].split()
        data_rows = [line.split() for line in table_content[1:]]
        # 转成DataFrame并指定数据类型为float
        step_df = pd.DataFrame(data_rows, columns=header, dtype=float)
        step_df["time_step"] = step_idx
        all_time_step_dfs.append(step_df)
    # 合并所有时间步的DataFrame
    combined_df = pd.concat(all_time_step_dfs, ignore_index=True)
    return combined_df

小提示

实际使用时一定要根据自己的仿真文件格式调整标记和过滤规则!不同仿真工具的verbose输出差异很大——有的用连续的=====做分隔,有的表头是固定的列名行,这些细节都得针对性修改才能准确提取表格内容。

内容的提问来源于stack exchange,提问作者Alemanio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:44:57