从文本文件提取CONTOUR数据导入pandas DataFrame的实现方案
需求说明
我有多个格式如下的文本文件,每个OBJECT块中以CONTOUR开头的行数量不固定,需要针对每个以OBJECT开头的区块,提取所有以CONTOUR开头的行内容,导入到pandas dataframe中,最终得到表头为Object、Points、Length的三列数据框。
文件格式示例
OBJECT 1 NAME: MT1(SP1) 3 contours object uses open contours. color (red, green, blue) = (0, 1, 0) CONTOUR #1,1,0 5 points length = 3.07e+006 pm CONTOUR #2,1,0 6 points length = 3.51e+006 pm CONTOUR #3,1,0 5 points length = 3.50e+006 pm OBJECT 2 NAME: MT2(SP3) 4 contours object uses open contours. color (red, green, blue) = (0, 1, 1) CONTOUR #1,2,0 4 points length = 1.86e+006 pm CONTOUR #2,2,0 4 points length = 2.29e+006 pm CONTOUR #3,2,0 5 points length = 2.47e+006 pm CONTOUR #3,2,0 5 points length = 2.47e+006 pm OBJECT 3 NAME: MT3(SP2) 1 contours object uses open contours. color (red, green, blue) = (1, 0, 1) CONTOUR #1,3,0 6 points length = 2.74e+006 pm
预期输出结果
Object | Points | Length 1 | 5 | 3.07e+006 1 | 6 | 3.51e+006 1 | 5 | 3.50e+006 2 | 4 | 1.86e+006 2 | 4 | 2.29e+006 2 | 5 | 2.47e+006 2 | 5 | 2.47e+006 3 | 6 | 2.74e+006
已尝试的方案
方法1:逐行判断标记提取区块
代码如下:
with open('textfile.txt', 'r') as input, open('new_textfile.txt', 'w') as output: for line in input: if line.strip() == "OBJECT 1": copy = True continue elif line.strip() == "OBJECT 2": copy = False continue elif copy: output.write(line) input.close() output.close()
该方法只能提取两个指定OBJECT之间的区块,无法适配数量不固定的OBJECT,也无法读取最后一个OBJECT区块(无后续OBJECT作为结束标记)。
方法2:正则匹配区块
代码如下:
pattern = 'OBJECT\s\d.*OBJECT\s\d' match = re.findall(pattern, text, re.DOTALL)
该方法存在区块边界识别问题,无法直接将提取内容转换为目标格式的pandas dataframe。
可行实现方案
实现思路
不需要提前分割OBJECT区块,逐行遍历文件时动态记录当前所属的OBJECT编号即可:
- 初始化变量存储当前OBJECT编号、存储结果的列表
- 逐行读取文件内容,遇到OBJECT开头的行就更新当前OBJECT编号
- 遇到CONTOUR开头的行,用正则提取Points和Length字段,和当前OBJECT编号一起存入结果列表
- 遍历完成后直接将结果列表转换为pandas dataframe
完整代码
import re import pandas as pd # 初始化存储数据的列表和当前OBJECT编号变量 data_list = [] current_object = None # 逐行读取文件 with open('test_sample.txt', 'r', encoding='utf-8') as f: for line in f: stripped_line = line.strip() # 匹配OBJECT行,更新当前编号 if stripped_line.startswith('OBJECT'): current_object = int(stripped_line.split()[1]) continue # 匹配CONTOUR行,提取所需字段 if stripped_line.startswith('CONTOUR'): # 提取Points值 points = int(re.search(r'(\d+)\s+points', stripped_line).group(1)) # 提取Length值 length = float(re.search(r'length\s*=\s*([\d.eE+-]+)', stripped_line).group(1)) data_list.append([current_object, points, length]) # 转换为DataFrame df = pd.DataFrame(data_list, columns=['Object', 'Points', 'Length']) # 如果需要输出为指定分隔符的文本,可以执行以下代码 # df.to_csv('output_result.txt', sep='|', index=False)
内容的提问来源于stack exchange,提问作者Christer Edvardsson
相关产品推荐
相关产品推荐

