如何编写程序提取文件中每行的Cell number与DCR对应数值
提取Cell序号与DCR数值的实现方案
针对你给出的文本格式,不管中间字段如何变化,只要每行保留Cell number:X和DCR:Y的格式,以下几种方法可以稳定提取目标信息:
Python 实现
用正则表达式匹配目标字段是最稳妥的方式,不受中间字段数量或顺序变化的影响:
import re # 读取文件内容(假设文件名为cell_data.txt) with open('cell_data.txt', 'r') as f: for line in f: line = line.strip() if not line: continue # 匹配Cell number后的数字和DCR后的数值 match = re.search(r'Cell number:(\d+),.*DCR:([\d.]+)', line) if match: cell_num = match.group(1) dcr_value = match.group(2) print(f"Cell {cell_num}: DCR = {dcr_value}")
解释:正则表达式Cell number:(\d+),.*DCR:([\d.]+)会精准定位Cell number:后的整数,以及DCR:后的小数,中间的.*会忽略任意内容,完全不受中间字段变化影响。
Shell(awk)实现
如果用命令行处理,awk可以快速完成提取:
awk 'match($0, /Cell number:([0-9]+).*DCR:([0-9.]+)/, arr) { print "Cell " arr[1] ": DCR = " arr[2] }' cell_data.txt
或者按字段遍历的写法(适合需要处理更多字段的场景):
awk -F, '{ for(i=1;i<=NF;i++){ if($i ~ /^Cell number:/) { split($i, arr, ":"); cell=arr[2] } if($i ~ /^DCR:/) { split($i, arr, ":"); dcr=arr[2] } } print "Cell " cell ": DCR = " dcr }' cell_data.txt
解释:第一种写法直接用正则匹配整行提取目标分组;第二种按逗号分割每行字段,遍历找到目标字段后拆分取值,两种方式都不受中间内容干扰。
内容的提问来源于stack exchange,提问作者Joao
相关产品推荐
相关产品推荐

