Python使用正则提取含重复Data-Information分段的实验数据及pandas处理咨询
解决方案
正则修改思路
- 原正则的中间通用匹配规则
((?:\r?\n.+)+)过于宽泛,无法准确定位Points区段,改为分层匹配逻辑更稳定:先拆分每个独立的Data-Information块,再在块内分别匹配头部信息和所有Points区段 - 每个Points区段的终止判定规则为:遇到下一个
Section:开头的行、遇到下一个Data-Information开头的行,或者到文件末尾
完整实现代码
无需依赖第三方库,直接用Python标准库+Pandas即可完成提取和格式化:
import re import pandas as pd # 读取原始数据文件 with open("你的数据文件路径.txt", "r", encoding="utf-8") as f: content = f.read() # 规则1:匹配所有独立的Data-Information块 info_pattern = re.compile( r"^Data-Information.*?(?=^Data-Information|\Z)", re.MULTILINE | re.DOTALL ) # 规则2:匹配每个块内的头部基础信息 head_pattern = re.compile( r"^Name:\s+(?P<Name>.+?)\s*$.*?^Sample:\s+(?P<Sample>.+?)\s*$.*?^System:\s+(?P<System>.+?)\s*$", re.MULTILINE | re.DOTALL ) # 规则3:匹配每个块内所有Points区段 point_pattern = re.compile( r"^Points\s+Time.*?(?=^Section:|\Z)", re.MULTILINE | re.DOTALL ) # 解析单个Points区段为DataFrame def parse_point_block(block: str, base_info: dict) -> pd.DataFrame: # 过滤空行,去除每行首尾空白 lines = [line.strip() for line in block.splitlines() if line.strip()] # 第一行为表头,第二行为单位行跳过,第三行开始为有效数据 header = lines[0].split() rows = [] for line in lines[2:]: # 处理带*** 1 ***格式的特殊行 line = line.replace("***", "").strip() row = line.split() if len(row) == len(header): rows.append(row) df = pd.DataFrame(rows, columns=header) # 附加当前块的基础信息,方便后续关联分析 for k, v in base_info.items(): df[k] = v return df # 批量处理所有数据 all_dfs = [] for info_block in info_pattern.findall(content): head_match = head_pattern.search(info_block) if not head_match: continue base_info = head_match.groupdict() # 提取当前块所有Points区段 for point_block in point_pattern.findall(info_block): df = parse_point_block(point_block, base_info) all_dfs.append(df) # 合并为全量数据表 final_df = pd.concat(all_dfs, ignore_index=True) # 可直接导出为csv或者做后续分析 # final_df.to_csv("提取结果.csv", index=False, encoding="utf-8-sig")
单个正则捕获方案(可选)
如果需要用单个正则一次性完成所有匹配,可使用支持重复分组捕获的regex第三方库,正则写法如下:
^Data-Information.*? ^Name:\s+(?P<Name>.+?)\s*$ ^Sample:\s+(?P<Sample>.+?)\s*$ .*? ^System:\s+(?P<System>.+?)\s*$ (?:.*?(?P<PointBlock>^Points\s+Time.*?(?=^Section:|^Data-Information|\Z)))*
匹配后调用captures("PointBlock")方法即可直接拿到当前Data-Information块下所有Points区段的内容。
内容的提问来源于stack exchange,提问作者Operator23
相关产品推荐
相关产品推荐

