You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shell脚本实现文件值映射去重写入的问题求助

解决Shell脚本输出内容重复问题

问题背景

现有两个文本文件:

  • projectnameid.txt:第一列为项目名,第二列为项目ID
  • endfilter3-all-b.txt:第一列为流水线ID,第二列为项目ID,第三列为分支名,第四列为时间戳

需求是读取endfilter3-all-b.txt,将其中的项目ID与projectnameid.txt匹配,把对应的项目名和ID写入finalized-project-names.txt。但编写的Shell脚本输出内容重复,需要修复。

相关文件内容

projectnameid.txt

basket-items                      9189
cloudfabric-notifications         10789
cloud-ports                       10965
common                            9081
customer-port-management          8550
deploy-quote                      8348
geographical-location-management  8549
internet-connections              9293
ipaddress                         8536
ip-addresses                      9294
order-notifications               11725
order-status                      8353
port-inventory                    8486
port-locations                    8490
pricing-quotes                    8493
product-availability              8488
product-catalogue                 8489
product-countrylist               8492
stub-service                      8510
customer-port-management-sf       10488
internet-connections-order-sf     11166
ip-addresses-order-sf             11165

endfilter3-all-b.txt

337718  10965  "refs/merge-requests/13/head"  "2023-07-19T11:39:41.739Z"
318933  8536   "develop"                      "2023-07-05T11:41:28.482Z"
366210  8549   "develop"                      "2023-08-11T13:49:18.905Z"
338835  8510   "main"                         "2023-07-20T06:45:59.823Z"
135208  8348   "main"                         "2023-02-17T11:25:07.723Z"
115402  8493   "main"                         "2023-02-07T06:52:05.486Z"
361979  9293   "refs/merge-requests/83/head"  "2023-08-09T07:38:32.831Z"
345703  11725  "main"                         "2023-07-26T08:31:11.004Z"
101775  8353   "main"                         "2023-02-02T09:22:47.402Z"
115414  8486   "main"                         "2023-02-07T07:41:35.478Z"
150861  9081   "main"                         "2023-03-13T05:37:31.370Z"
135733  8489   "main"                         "2023-02-17T16:14:51.280Z"

预期输出finalized-project-names.txt

cloud-ports                 10965
ipaddress               8536
geographical-location-management    8549
stub-service                8510 
deploy-quote                8348   
pricing-quotes              8493   
internet-connections            9293   
order-notifications         11725  
order-status                8353   
port-inventory              8486   
common                  9081   
product-catalogue           8489    

尝试的脚本

#!/bin/bash

> finalized-project-names.txt

project_name_id_file="projectnameid.txt"
while read -r project_name project_id; do
   project_name_id_map[$project_id]=$project_name
done < "$project_name_id_file"

endfilter3_all_b_file="endfilter3-all-b.txt"
while read pipeline_id project_id branch_name timestamp; do
  project_name=${project_name_id_map[$project_id]}
  if [ -n "$project_name" ]; then
    echo "$project_name $project_id" >> finalized-project-names.txt
  fi
done < "$endfilter3_all_b_file"

实际错误输出

cloud-ports 10965
ipaddress 8536
geographical-location-management 8549
stub-service 8510
deploy-quote 8348
pricing-quotes 8493
internet-connections 9293
order-notifications 11725
order-status 8353
port-inventory 8486
common 9081
product-catalogue 8489
cloud-ports 10965
ipaddress 8536
geographical-location-management 8549
stub-service 8510
deploy-quote 8348
pricing-quotes 8493
internet-connections 9293
order-notifications 11725
order-status 8353
port-inventory 8486
common 9081
product-catalogue 8489

问题原因与修复方案

问题原因

你的脚本逻辑本身是正确的,输出重复的核心原因是脚本被执行了两次,且第二次执行时可能误将脚本输出追加到了已有的文件中(比如用./your_script.sh >> finalized-project-names.txt的方式运行),导致内容重复写入。

修复方案

方案1:规范脚本运行方式

脚本开头已经通过> finalized-project-names.txt清空了输出文件,只需直接运行脚本(./your_script.sh),不要额外追加输出,就能得到正确的12行结果。

方案2:添加去重逻辑

如果担心后续可能出现重复内容,可以在脚本末尾添加去重步骤,确保每个项目只保留一条记录:

#!/bin/bash

# 清空输出文件
> finalized-project-names.txt

# 加载项目名-ID映射
declare -A project_name_id_map
project_name_id_file="projectnameid.txt"
while read -r project_name project_id; do
   project_name_id_map[$project_id]=$project_name
done < "$project_name_id_file"

# 匹配并写入结果
endfilter3_all_b_file="endfilter3-all-b.txt"
while read -r pipeline_id project_id branch_name timestamp; do
  project_name=${project_name_id_map[$project_id]}
  if [ -n "$project_name" ]; then
    echo "$project_name $project_id" >> finalized-project-names.txt
  fi
done < "$endfilter3_all_b_file"

# 去重并保留原顺序
awk '!seen[$0]++' finalized-project-names.txt > temp.txt && mv temp.txt finalized-project-names.txt

注:awk '!seen[$0]++'会按原顺序去重,不会改变输出的排列顺序。

方案3:用awk一步完成(更简洁高效)

直接使用awk处理可以避免Shell循环的潜在问题,同时天然保证每个匹配项只处理一次:

awk '
    NR==FNR { map[$2]=$1; next }
    $2 in map { print map[$2], $2 }
' projectnameid.txt endfilter3-all-b.txt > finalized-project-names.txt

逻辑说明:

  1. 先读取projectnameid.txt,建立项目ID到项目名的映射表map
  2. 再读取endfilter3-all-b.txt,如果第二列(项目ID)存在于映射表中,就输出对应的项目名和ID
  3. 最后直接重定向到输出文件,自动清空并写入,彻底避免重复问题

内容的提问来源于stack exchange,提问作者MisterJay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:02:02