如何用Bash修改制表符文件中符合条件的第二列内容
问题描述
我有一个制表符分隔的txt文件,内容如下:
Serving Sector Target Sector HO Attempts HO Successful Attempts 1002080 1002081 8 8 1002080 1002084 0 0 1002080 1002974 2 2 1002080 2104-2975 5 5 1002080 1002976 2 2 1002080 1012237 10 10 1002080 1012281 0 0
部分Target Sector(第二列)的格式为ABCD-YYYY(如2104-2975),需将其转换为BC0YYYY格式(即1002975)。
我已编写如下Bash脚本片段:
while read -r line; do if echo $line | grep -E '([0-9])-([0-9])' # If line matches criteria then string=`echo "$line" | awk -F '\t' '{print $2}'` #fetch column 2 LAC=${string%-*} #LAC= ABCD CI=${string##*-} #CI = YYYY if [ ${#CI} -lt 5 ]; then CI="0"$CI; #IF stringlength of CI is less than 5, add 0 fi LAC2=`echo $LAC | cut -c2-3` #LAC2 = BC GERANCELL=$LAC2$CI fi done < input.txt
请问如何将该行的第二列更新为新值$GERANCELL?
解决方案
你可以通过数组分割行内容直接替换第二列,同时优化脚本效率(减少不必要的子进程调用)。修改后的完整脚本如下:
while read -r -a fields; do # 表头直接输出,不处理 if [[ "${fields[0]}" == "Serving" ]]; then echo -e "${fields[0]}\t${fields[1]}\t${fields[2]}\t${fields[3]}" continue fi target_sector="${fields[1]}" # 判断第二列是否带分隔符'-' if [[ "$target_sector" == *-* ]]; then LAC="${target_sector%-*}" CI="${target_sector##*-}" # CI不足5位补0 if [[ ${#CI} -lt 5 ]]; then CI="0$CI" fi # 提取LAC的第2到第3位 LAC2="${LAC:1:2}" # 拼接成目标格式BC0YYYY GERANCELL="${LAC2}0${CI}" # 替换第二列后输出整行 echo -e "${fields[0]}\t${GERANCELL}\t${fields[2]}\t${fields[3]}" else # 不需要转换的行直接输出 echo -e "${fields[0]}\t${fields[1]}\t${fields[2]}\t${fields[3]}" fi done < input.txt > output.txt
核心优化点:
- 用
read -a fields把每行按制表符拆成数组,无需调用awk提取列,提升运行效率 - 增加表头判断,避免表头被错误处理
- 用bash内置的字符串切片
${LAC:1:2}替代cut命令,减少子进程调用 - 直接构造新行输出,重定向到
output.txt避免覆盖原文件 - 明确拼接时加入
0,严格符合BC0YYYY的格式要求
如果你的文件较大,用纯awk实现会更高效,代码也更简洁:
BEGIN { FS="\t"; OFS="\t" } # 第一行表头直接输出 NR==1 { print; next } # 匹配第二列带'-'的行 $2 ~ /[0-9]+-[0-9]+/ { split($2, arr, "-") lac = arr[1] ci = arr[2] if (length(ci) < 5) ci = "0" ci lac2 = substr(lac, 2, 2) $2 = lac2 "0" ci } # 输出所有行 1
使用方式:
awk -f convert.awk input.txt > output.txt
内容的提问来源于stack exchange,提问作者SHR
相关产品推荐
相关产品推荐

