基于Error字段条件的awk文件解析与更新技术问询
制表符分隔文件处理需求
- 处理规则:
- 若第7字段(Error字段)包含new或update,需提取其中
aaa+数字格式的编号填入第5字段(ID2),并将第6字段(Status)设为"New" - 若第7字段不含上述关键词,则跳过更新,保留原字段内容(空字段保持为空)
- 若第7字段(Error字段)包含new或update,需提取其中
- 数据结构(制表符分隔,字段顺序):#、xxx、ID、Pos、ID2、Status、Error、Comment
输入文件示例
#,xxx,ID,Pos,ID2,Status,Error,Comment 135,,ABCD1,153740155,,,[Cells: AH8] .... 135,,ACADVL,7222674,,,[Cells: F9-J9] location (1) is inconsistent with (3) 135,,ACP5,11575174,,,"[Cells: D11-T11, AB11-AD11] This record is new and update, submitted aaa000000001 already.,aaa 136,,ACTA2,88943843,,,"[Cells: D12-T12, AB12-AD12] This record is new and update, submitted aaa000000002 already. 136,,ACVRL1,51913357,,,[Cells: F140-J140] location (1) is inconsistent with (3),aaa
期望输出示例
#,xxx,ID,Pos,ID2,Status,Error,Comment 135,,ABCD1,153740155,,,[Cells: AH8] .... 135,,ACADVL,7222674,,,[Cells: F9-J9] location (1) is inconsistent with (3) 135,,ACP5,11575174,aaa000000001,New,"[Cells: D11-T11, AB11-AD11] This record is new and update, submitted aaa000000001 already.,aaa 136,,ACTA2,88943843,aaa000000002,New,"[Cells: D12-T12, AB12-AD12] This record is new and update, submitted aaa000000002 already. 136,,ACVRL1,51913357,,,[Cells: F140-J140] location (1) is inconsistent with (3),aaa
原脚本问题分析
原awk脚本存在以下缺陷:
- 未判断Error字段是否包含new/update,直接处理所有行
- 正则
/aaa[0-9]/仅匹配单个数字,无法提取完整的aaa+连续数字编号 - 错误修改了第7字段(Error)的内容,需求要求保留原Error字段
- 循环逻辑冗余,无需循环即可提取目标编号
- Status字段值应为"New"(首字母大写),原脚本写为"new"
修正后的awk脚本
BEGIN { FS = "\t"; OFS = "\t"; } # 表头行直接输出,不做处理 NR == 1 { print $0; next; } { # 判断Error字段是否包含new或update if ($7 ~ /new|update/) { # 提取aaa后跟连续数字的完整匹配项 match($7, /aaa[0-9]+/, arr); # 确保提取到有效编号后再更新字段 if (arr[0] != "") { $5 = arr[0]; $6 = "New"; } } # 输出处理后的行(未匹配条件的行保持原样) print $0; }
脚本说明
- 表头行直接打印,避免处理逻辑干扰
- 使用
$7 ~ /new|update/精准匹配需要处理的行 - 用
match()函数结合正则/aaa[0-9]+/提取完整的目标编号 - 仅在成功提取到编号时更新ID2和Status字段,防止空值覆盖
- 完全保留原Error字段内容,符合需求要求
内容的提问来源于stack exchange,提问作者justaguy
相关产品推荐
相关产品推荐

