You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历文本文件提取RowID及关联错误警告并构建DataFrame

问题需求

需要遍历一份文本报告文件,提取**RowID(仅保留整数)**以及对应Name行下方的所有错误/警告信息,忽略其他内容,最终将提取结果组装成Pandas DataFrame。

示例输入文本

***********************************************************
Date/Time:    6/5/2023
FileName:    somefile.txt
Report:        Standard Report
***********************************************************
Success:    1234
Failures:    1234

RowID:    100
Name:    Smith, John
    This person did not meet the criteria because of x,y,z.
RowID:    101
Name:    Smith, Susie
    This is a warning.
    This is an error.
    Criteria was not met.
RowID:    103
Name:    Jones, Bob
    This person had invalid characters in the email field.

现有尝试代码

search_string = "RowID:"
next_search_string = "Name:"

with open('report.txt') as y:
    for line in y:
        if line.startswith(search_string):
            print(line.split(':')[1].strip())
        if line.startswith(next_search_string):
            print(next(y))
            while not (next(y)).startswith(search_string):
                print(next(y))
        if (next(y)).startswith(search_string):
            pass

期望输出格式

100, This person did not meet the criteria because of x,y,z. 
101, This is a warning.  This is an error.  Criteria was not met. 
103, This person had invalid characters in the email field.

正确实现方案

思路说明

  1. 用索引遍历文件行,避免next()导致的行跳过问题,逻辑更可控;
  2. 捕获RowID时直接转换为整数,确保数据类型准确;
  3. 遇到Name行后,持续读取后续行直到下一个RowID出现或文件结束,收集所有非空的错误/警告文本;
  4. 将收集到的RowID和合并后的文本整理为列表,最终转换为DataFrame。

完整代码

import pandas as pd

def parse_report(file_path):
    data = []
    current_rowid = None
    current_messages = []
    
    with open(file_path, 'r') as f:
        lines = f.readlines()
        idx = 0
        total_lines = len(lines)
        
        while idx < total_lines:
            line = lines[idx].strip()
            # 提取RowID整数
            if line.startswith('RowID:'):
                current_rowid = int(line.split(':')[1].strip())
                idx += 1
            # 遇到Name行后开始收集消息
            elif line.startswith('Name:'):
                idx += 1
                # 循环读取直到下一个RowID或文件结束
                while idx < total_lines:
                    msg_line = lines[idx].strip()
                    if msg_line.startswith('RowID:'):
                        break
                    # 跳过空行,仅收集有效消息
                    if msg_line:
                        current_messages.append(msg_line)
                    idx += 1
                # 将当前RowID和合并后的消息存入数据列表
                if current_rowid is not None and current_messages:
                    combined_msg = ' '.join(current_messages)
                    data.append({'RowID': current_rowid, 'Messages': combined_msg})
                    # 重置变量准备下一轮
                    current_rowid = None
                    current_messages = []
            else:
                idx += 1
    
    # 转换为DataFrame
    return pd.DataFrame(data)

# 调用示例
df = parse_report('report.txt')

# 打印匹配期望格式的结果
for _, row in df.iterrows():
    print(f"{row['RowID']}, {row['Messages']}")

# 查看生成的DataFrame
print("\n生成的DataFrame:")
print(df)

代码输出示例

运行后会先输出:

100, This person did not meet the criteria because of x,y,z.
101, This is a warning. This is an error. Criteria was not met.
103, This person had invalid characters in the email field.

然后输出DataFrame:

生成的DataFrame:
   RowID                                           Messages
0    100  This person did not meet the criteria because ...
1    101  This is a warning. This is an error. Criteria ...
2    103  This person had invalid characters in the emai...

内容的提问来源于stack exchange,提问作者ByRequest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:50:02