遍历文本文件提取RowID及关联错误警告并构建DataFrame
问题需求
需要遍历一份文本报告文件,提取**RowID(仅保留整数)**以及对应Name行下方的所有错误/警告信息,忽略其他内容,最终将提取结果组装成Pandas DataFrame。
示例输入文本
*********************************************************** Date/Time: 6/5/2023 FileName: somefile.txt Report: Standard Report *********************************************************** Success: 1234 Failures: 1234 RowID: 100 Name: Smith, John This person did not meet the criteria because of x,y,z. RowID: 101 Name: Smith, Susie This is a warning. This is an error. Criteria was not met. RowID: 103 Name: Jones, Bob This person had invalid characters in the email field.
现有尝试代码
search_string = "RowID:" next_search_string = "Name:" with open('report.txt') as y: for line in y: if line.startswith(search_string): print(line.split(':')[1].strip()) if line.startswith(next_search_string): print(next(y)) while not (next(y)).startswith(search_string): print(next(y)) if (next(y)).startswith(search_string): pass
期望输出格式
100, This person did not meet the criteria because of x,y,z. 101, This is a warning. This is an error. Criteria was not met. 103, This person had invalid characters in the email field.
正确实现方案
思路说明
- 用索引遍历文件行,避免
next()导致的行跳过问题,逻辑更可控; - 捕获RowID时直接转换为整数,确保数据类型准确;
- 遇到Name行后,持续读取后续行直到下一个RowID出现或文件结束,收集所有非空的错误/警告文本;
- 将收集到的RowID和合并后的文本整理为列表,最终转换为DataFrame。
完整代码
import pandas as pd def parse_report(file_path): data = [] current_rowid = None current_messages = [] with open(file_path, 'r') as f: lines = f.readlines() idx = 0 total_lines = len(lines) while idx < total_lines: line = lines[idx].strip() # 提取RowID整数 if line.startswith('RowID:'): current_rowid = int(line.split(':')[1].strip()) idx += 1 # 遇到Name行后开始收集消息 elif line.startswith('Name:'): idx += 1 # 循环读取直到下一个RowID或文件结束 while idx < total_lines: msg_line = lines[idx].strip() if msg_line.startswith('RowID:'): break # 跳过空行,仅收集有效消息 if msg_line: current_messages.append(msg_line) idx += 1 # 将当前RowID和合并后的消息存入数据列表 if current_rowid is not None and current_messages: combined_msg = ' '.join(current_messages) data.append({'RowID': current_rowid, 'Messages': combined_msg}) # 重置变量准备下一轮 current_rowid = None current_messages = [] else: idx += 1 # 转换为DataFrame return pd.DataFrame(data) # 调用示例 df = parse_report('report.txt') # 打印匹配期望格式的结果 for _, row in df.iterrows(): print(f"{row['RowID']}, {row['Messages']}") # 查看生成的DataFrame print("\n生成的DataFrame:") print(df)
代码输出示例
运行后会先输出:
100, This person did not meet the criteria because of x,y,z. 101, This is a warning. This is an error. Criteria was not met. 103, This person had invalid characters in the email field.
然后输出DataFrame:
生成的DataFrame: RowID Messages 0 100 This person did not meet the criteria because ... 1 101 This is a warning. This is an error. Criteria ... 2 103 This person had invalid characters in the emai...
内容的提问来源于stack exchange,提问作者ByRequest
相关产品推荐
相关产品推荐

