如何将多帧协议转储文本文件转换为Pandas DataFrame?
解决方案
直接用Python结合Pandas处理即可,无需复杂的多行正则,按帧分组处理是最简洁的方式:
核心思路
- 分割帧数据:以
Frame开头的行为标记,把整个文本分割成独立的帧块 - 提取关键信息:从每个帧块中提取帧编号,将帧内所有内容作为详情文本
- 构造并格式化显示:生成DataFrame后,自定义打印逻辑模拟需求中的表格样式
完整代码示例
import pandas as pd import re # 读取原始协议转储文件 with open('hci_dump.txt', 'r', encoding='utf-8') as f: raw_content = f.read() # 分割出所有独立帧(正向预查确保每个帧块包含完整的Frame开头行) frame_blocks = [block.strip() for block in re.split(r'(?=Frame \d+:)', raw_content) if block.strip()] # 整理成DataFrame所需的结构化数据 frame_data = [] for block in frame_blocks: lines = block.split('\n') # 提取帧编号 num_match = re.search(r'Frame (\d+):', lines[0]) if num_match: frame_number = num_match.group(1) # 保留帧内所有原始行作为详情 frame_details = '\n'.join(lines) frame_data.append({'FrameNumber': frame_number, 'Details': frame_details}) # 创建DataFrame df = pd.DataFrame(frame_data) # 自定义打印函数,输出需求中的表格样式 def print_formatted_table(df): separator_line = '|' + '-'*79 + '|' # 打印表头 print(f" FrameNumber Details ") print(separator_line) # 遍历每个帧的内容 for idx, row in df.iterrows(): detail_lines = row['Details'].split('\n') # 逐行打印,仅在第三行(Direction行)显示帧编号 for line_idx, line in enumerate(detail_lines): if line_idx == 2: print(f"| {row['FrameNumber']:4} | {line.ljust(75)} |") else: print(f"| | {line.ljust(75)} |") # 打印分隔线,最后一行用+结尾 print(separator_line if idx != len(df)-1 else '+' + '-'*79 + '+') # 执行打印 print_formatted_table(df)
说明
- 替换代码中的
hci_dump.txt为你的实际文件路径 - 正则
(?=Frame \d+:)用于精准分割帧,不会破坏帧的开头行 - 自定义打印函数严格匹配你需要的表格样式,帧编号仅显示在
[Direction:]对应的行 - 生成的
df是标准的Pandas DataFrame,可用于后续的数据分析或导出
内容的提问来源于stack exchange,提问作者pixelworks
相关产品推荐
相关产品推荐

