You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Unix Shell脚本:竖线分隔文件指定列的字符串过滤需求

处理竖线分隔文件指定列的Shell脚本方案

我来给你提供一个基于awk的Shell脚本实现,它能精准处理你描述的需求:只修改指定列,保留匹配IPL_前缀的字符串,清理无关内容和多余的分号,其他列保持原样。

核心思路

  1. 以竖线|作为输入输出的分隔符,确保列结构不变
  2. 跳过表头行直接输出
  3. 对目标列先按逗号分割成多个子块(比如CRIC1:xxx,CRIC2:yyy)
  4. 每个子块再拆分前缀(如CRIC1)和项目列表(如IPL_M1;TEST_M1)
  5. 筛选项目列表中以IPL_开头的项,用分号重新拼接
  6. 将处理后的子块用逗号拼接,替换原目标列后输出整行

脚本实现

针对简化场景(两列,处理第2列)

将以下内容保存为process_simple.awk:

BEGIN { FS = OFS = "|" }
# 输出表头
NR == 1 { print; next }
{
    # 按逗号分割第2列的子块
    split($2, parts, ",")
    new_col = ""
    for (i in parts) {
        # 拆分前缀和项目列表
        split(parts[i], kv_pair, ":")
        prefix = kv_pair[1]
        split(kv_pair[2], items, ";")
        
        filtered_items = ""
        # 筛选IPL_开头的项目
        for (j in items) {
            if (items[j] ~ /^IPL_/) {
                if (filtered_items != "") filtered_items = filtered_items ";"
                filtered_items = filtered_items items[j]
            }
        }
        # 拼接子块到新列
        if (new_col != "") new_col = new_col ","
        new_col = new_col prefix ":" filtered_items
    }
    # 替换原列并输出
    $2 = new_col
    print
}

运行命令:

awk -f process_simple.awk input_simple.txt > output_simple.txt

针对5列场景(处理第4列)

将以下内容保存为process_full.awk:

BEGIN { FS = OFS = "|" }
# 输出表头
NR == 1 { print; next }
{
    # 按逗号分割第4列的子块
    split($4, parts, ",")
    new_col = ""
    for (i in parts) {
        # 拆分前缀和项目列表
        split(parts[i], kv_pair, ":")
        prefix = kv_pair[1]
        split(kv_pair[2], items, ";")
        
        filtered_items = ""
        # 筛选IPL_开头的项目
        for (j in items) {
            if (items[j] ~ /^IPL_/) {
                if (filtered_items != "") filtered_items = filtered_items ";"
                filtered_items = filtered_items items[j]
            }
        }
        # 拼接子块到新列
        if (new_col != "") new_col = new_col ","
        new_col = new_col prefix ":" filtered_items
    }
    # 替换原列并输出
    $4 = new_col
    print
}

运行命令:

awk -f process_full.awk input_full.txt > output_full.txt

效果验证

以你提供的5列示例为例,输入文件经过脚本处理后,会完全匹配你给出的输出结果:

  • 第4列中所有非IPL_开头的项目会被移除
  • 多余的分号会自动清理(比如CRIC1:后面不会留下空分号)
  • 无匹配内容的子块会保留前缀加空值(如CRIC1:)
  • 其他列内容完全不变

内容的提问来源于stack exchange,提问作者Saquib Hussain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:59:52