You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从文本文件提取CONTOUR数据导入pandas DataFrame的实现方案

需求说明

我有多个格式如下的文本文件,每个OBJECT块中以CONTOUR开头的行数量不固定,需要针对每个以OBJECT开头的区块,提取所有以CONTOUR开头的行内容,导入到pandas dataframe中,最终得到表头为Object、Points、Length的三列数据框。

文件格式示例

OBJECT 1
NAME:  MT1(SP1)
       3 contours
       object uses open contours.
       color (red, green, blue) = (0, 1, 0)

    CONTOUR #1,1,0  5 points    length = 3.07e+006 pm
    CONTOUR #2,1,0  6 points    length = 3.51e+006 pm
    CONTOUR #3,1,0  5 points    length = 3.50e+006 pm

OBJECT 2
NAME:  MT2(SP3)
       4 contours
       object uses open contours.
       color (red, green, blue) = (0, 1, 1)

    CONTOUR #1,2,0  4 points    length = 1.86e+006 pm
    CONTOUR #2,2,0  4 points    length = 2.29e+006 pm
    CONTOUR #3,2,0  5 points    length = 2.47e+006 pm
    CONTOUR #3,2,0  5 points    length = 2.47e+006 pm

OBJECT 3
NAME:  MT3(SP2)
       1 contours
       object uses open contours.
       color (red, green, blue) = (1, 0, 1)

    CONTOUR #1,3,0  6 points    length = 2.74e+006 pm

预期输出结果

Object | Points | Length
1 | 5 | 3.07e+006   
1 | 6 | 3.51e+006
1 | 5 | 3.50e+006
2 | 4 | 1.86e+006
2 | 4 | 2.29e+006
2 | 5 | 2.47e+006
2 | 5 | 2.47e+006
3 | 6 | 2.74e+006

已尝试的方案

方法1:逐行判断标记提取区块

代码如下:

with open('textfile.txt', 'r') as input, open('new_textfile.txt', 'w') as output:
  for line in input:
        if line.strip() == "OBJECT 1":
            copy = True
            continue
        elif line.strip() == "OBJECT 2":
            copy = False
            continue
        elif copy:
            output.write(line)
            
input.close()
output.close()

该方法只能提取两个指定OBJECT之间的区块,无法适配数量不固定的OBJECT,也无法读取最后一个OBJECT区块(无后续OBJECT作为结束标记)。

方法2:正则匹配区块

代码如下:

pattern = 'OBJECT\s\d.*OBJECT\s\d'
match = re.findall(pattern, text, re.DOTALL)

该方法存在区块边界识别问题,无法直接将提取内容转换为目标格式的pandas dataframe。

可行实现方案

实现思路

不需要提前分割OBJECT区块,逐行遍历文件时动态记录当前所属的OBJECT编号即可:

  1. 初始化变量存储当前OBJECT编号、存储结果的列表
  2. 逐行读取文件内容,遇到OBJECT开头的行就更新当前OBJECT编号
  3. 遇到CONTOUR开头的行,用正则提取Points和Length字段,和当前OBJECT编号一起存入结果列表
  4. 遍历完成后直接将结果列表转换为pandas dataframe

完整代码

import re
import pandas as pd

# 初始化存储数据的列表和当前OBJECT编号变量
data_list = []
current_object = None

# 逐行读取文件
with open('test_sample.txt', 'r', encoding='utf-8') as f:
    for line in f:
        stripped_line = line.strip()
        # 匹配OBJECT行,更新当前编号
        if stripped_line.startswith('OBJECT'):
            current_object = int(stripped_line.split()[1])
            continue
        # 匹配CONTOUR行,提取所需字段
        if stripped_line.startswith('CONTOUR'):
            # 提取Points值
            points = int(re.search(r'(\d+)\s+points', stripped_line).group(1))
            # 提取Length值
            length = float(re.search(r'length\s*=\s*([\d.eE+-]+)', stripped_line).group(1))
            data_list.append([current_object, points, length])

# 转换为DataFrame
df = pd.DataFrame(data_list, columns=['Object', 'Points', 'Length'])

# 如果需要输出为指定分隔符的文本,可以执行以下代码
# df.to_csv('output_result.txt', sep='|', index=False)

内容的提问来源于stack exchange,提问作者Christer Edvardsson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 22:45:08