You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用正则提取含重复Data-Information分段的实验数据及pandas处理咨询

解决方案

正则修改思路

  • 原正则的中间通用匹配规则((?:\r?\n.+)+)过于宽泛,无法准确定位Points区段,改为分层匹配逻辑更稳定:先拆分每个独立的Data-Information块,再在块内分别匹配头部信息和所有Points区段
  • 每个Points区段的终止判定规则为:遇到下一个Section:开头的行、遇到下一个Data-Information开头的行,或者到文件末尾

完整实现代码

无需依赖第三方库,直接用Python标准库+Pandas即可完成提取和格式化:

import re
import pandas as pd

# 读取原始数据文件
with open("你的数据文件路径.txt", "r", encoding="utf-8") as f:
    content = f.read()

# 规则1:匹配所有独立的Data-Information块
info_pattern = re.compile(
    r"^Data-Information.*?(?=^Data-Information|\Z)",
    re.MULTILINE | re.DOTALL
)
# 规则2:匹配每个块内的头部基础信息
head_pattern = re.compile(
    r"^Name:\s+(?P<Name>.+?)\s*$.*?^Sample:\s+(?P<Sample>.+?)\s*$.*?^System:\s+(?P<System>.+?)\s*$",
    re.MULTILINE | re.DOTALL
)
# 规则3:匹配每个块内所有Points区段
point_pattern = re.compile(
    r"^Points\s+Time.*?(?=^Section:|\Z)",
    re.MULTILINE | re.DOTALL
)

# 解析单个Points区段为DataFrame
def parse_point_block(block: str, base_info: dict) -> pd.DataFrame:
    # 过滤空行,去除每行首尾空白
    lines = [line.strip() for line in block.splitlines() if line.strip()]
    # 第一行为表头,第二行为单位行跳过,第三行开始为有效数据
    header = lines[0].split()
    rows = []
    for line in lines[2:]:
        # 处理带*** 1 ***格式的特殊行
        line = line.replace("***", "").strip()
        row = line.split()
        if len(row) == len(header):
            rows.append(row)
    df = pd.DataFrame(rows, columns=header)
    # 附加当前块的基础信息,方便后续关联分析
    for k, v in base_info.items():
        df[k] = v
    return df

# 批量处理所有数据
all_dfs = []
for info_block in info_pattern.findall(content):
    head_match = head_pattern.search(info_block)
    if not head_match:
        continue
    base_info = head_match.groupdict()
    # 提取当前块所有Points区段
    for point_block in point_pattern.findall(info_block):
        df = parse_point_block(point_block, base_info)
        all_dfs.append(df)

# 合并为全量数据表
final_df = pd.concat(all_dfs, ignore_index=True)
# 可直接导出为csv或者做后续分析
# final_df.to_csv("提取结果.csv", index=False, encoding="utf-8-sig")

单个正则捕获方案(可选)

如果需要用单个正则一次性完成所有匹配,可使用支持重复分组捕获的regex第三方库,正则写法如下:

^Data-Information.*?
^Name:\s+(?P<Name>.+?)\s*$
^Sample:\s+(?P<Sample>.+?)\s*$
.*?
^System:\s+(?P<System>.+?)\s*$
(?:.*?(?P<PointBlock>^Points\s+Time.*?(?=^Section:|^Data-Information|\Z)))*

匹配后调用captures("PointBlock")方法即可直接拿到当前Data-Information块下所有Points区段的内容。

内容的提问来源于stack exchange,提问作者Operator23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 22:15:06