You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Bash修改制表符文件中符合条件的第二列内容

问题描述

我有一个制表符分隔的txt文件,内容如下:

Serving Sector  Target Sector   HO Attempts HO Successful Attempts
1002080 1002081 8   8
1002080 1002084 0   0
1002080 1002974 2   2
1002080 2104-2975   5   5
1002080 1002976 2   2
1002080 1012237 10  10
1002080 1012281 0   0

部分Target Sector(第二列)的格式为ABCD-YYYY(如2104-2975),需将其转换为BC0YYYY格式(即1002975)。

我已编写如下Bash脚本片段:

while read -r line;
do
        if echo $line | grep -E '([0-9])-([0-9])' # If line matches criteria
        then
                string=`echo "$line" | awk -F '\t' '{print $2}'`   #fetch column 2
                LAC=${string%-*}   #LAC= ABCD
                CI=${string##*-}   #CI = YYYY
                if [ ${#CI} -lt 5 ]; then CI="0"$CI;  #IF stringlength of CI is less than 5, add 0 
                fi
                LAC2=`echo $LAC | cut -c2-3` #LAC2 = BC
                GERANCELL=$LAC2$CI

        fi
done < input.txt

请问如何将该行的第二列更新为新值$GERANCELL?

解决方案

你可以通过数组分割行内容直接替换第二列,同时优化脚本效率(减少不必要的子进程调用)。修改后的完整脚本如下:

while read -r -a fields; do
    # 表头直接输出,不处理
    if [[ "${fields[0]}" == "Serving" ]]; then
        echo -e "${fields[0]}\t${fields[1]}\t${fields[2]}\t${fields[3]}"
        continue
    fi

    target_sector="${fields[1]}"
    # 判断第二列是否带分隔符'-'
    if [[ "$target_sector" == *-* ]]; then
        LAC="${target_sector%-*}"
        CI="${target_sector##*-}"
        # CI不足5位补0
        if [[ ${#CI} -lt 5 ]]; then
            CI="0$CI"
        fi
        # 提取LAC的第2到第3位
        LAC2="${LAC:1:2}"
        # 拼接成目标格式BC0YYYY
        GERANCELL="${LAC2}0${CI}"
        
        # 替换第二列后输出整行
        echo -e "${fields[0]}\t${GERANCELL}\t${fields[2]}\t${fields[3]}"
    else
        # 不需要转换的行直接输出
        echo -e "${fields[0]}\t${fields[1]}\t${fields[2]}\t${fields[3]}"
    fi
done < input.txt > output.txt

核心优化点:

  • 用read -a fields把每行按制表符拆成数组,无需调用awk提取列,提升运行效率
  • 增加表头判断,避免表头被错误处理
  • 用bash内置的字符串切片${LAC:1:2}替代cut命令,减少子进程调用
  • 直接构造新行输出,重定向到output.txt避免覆盖原文件
  • 明确拼接时加入0,严格符合BC0YYYY的格式要求

如果你的文件较大,用纯awk实现会更高效,代码也更简洁:

BEGIN { FS="\t"; OFS="\t" }
# 第一行表头直接输出
NR==1 { print; next }
# 匹配第二列带'-'的行
$2 ~ /[0-9]+-[0-9]+/ {
    split($2, arr, "-")
    lac = arr[1]
    ci = arr[2]
    if (length(ci) < 5) ci = "0" ci
    lac2 = substr(lac, 2, 2)
    $2 = lac2 "0" ci
}
# 输出所有行
1

使用方式:

awk -f convert.awk input.txt > output.txt

内容的提问来源于stack exchange,提问作者SHR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 18:15:42