Unix Shell脚本:竖线分隔文件指定列的字符串过滤需求
处理竖线分隔文件指定列的Shell脚本方案
我来给你提供一个基于awk的Shell脚本实现,它能精准处理你描述的需求:只修改指定列,保留匹配IPL_前缀的字符串,清理无关内容和多余的分号,其他列保持原样。
核心思路
- 以竖线
|作为输入输出的分隔符,确保列结构不变 - 跳过表头行直接输出
- 对目标列先按逗号分割成多个子块(比如
CRIC1:xxx,CRIC2:yyy) - 每个子块再拆分前缀(如
CRIC1)和项目列表(如IPL_M1;TEST_M1) - 筛选项目列表中以
IPL_开头的项,用分号重新拼接 - 将处理后的子块用逗号拼接,替换原目标列后输出整行
脚本实现
针对简化场景(两列,处理第2列)
将以下内容保存为process_simple.awk:
BEGIN { FS = OFS = "|" } # 输出表头 NR == 1 { print; next } { # 按逗号分割第2列的子块 split($2, parts, ",") new_col = "" for (i in parts) { # 拆分前缀和项目列表 split(parts[i], kv_pair, ":") prefix = kv_pair[1] split(kv_pair[2], items, ";") filtered_items = "" # 筛选IPL_开头的项目 for (j in items) { if (items[j] ~ /^IPL_/) { if (filtered_items != "") filtered_items = filtered_items ";" filtered_items = filtered_items items[j] } } # 拼接子块到新列 if (new_col != "") new_col = new_col "," new_col = new_col prefix ":" filtered_items } # 替换原列并输出 $2 = new_col print }
运行命令:
awk -f process_simple.awk input_simple.txt > output_simple.txt
针对5列场景(处理第4列)
将以下内容保存为process_full.awk:
BEGIN { FS = OFS = "|" } # 输出表头 NR == 1 { print; next } { # 按逗号分割第4列的子块 split($4, parts, ",") new_col = "" for (i in parts) { # 拆分前缀和项目列表 split(parts[i], kv_pair, ":") prefix = kv_pair[1] split(kv_pair[2], items, ";") filtered_items = "" # 筛选IPL_开头的项目 for (j in items) { if (items[j] ~ /^IPL_/) { if (filtered_items != "") filtered_items = filtered_items ";" filtered_items = filtered_items items[j] } } # 拼接子块到新列 if (new_col != "") new_col = new_col "," new_col = new_col prefix ":" filtered_items } # 替换原列并输出 $4 = new_col print }
运行命令:
awk -f process_full.awk input_full.txt > output_full.txt
效果验证
以你提供的5列示例为例,输入文件经过脚本处理后,会完全匹配你给出的输出结果:
- 第4列中所有非
IPL_开头的项目会被移除 - 多余的分号会自动清理(比如
CRIC1:后面不会留下空分号) - 无匹配内容的子块会保留前缀加空值(如
CRIC1:) - 其他列内容完全不变
内容的提问来源于stack exchange,提问作者Saquib Hussain
相关产品推荐
相关产品推荐

