如何用bash或Python实现CSV与TXT列匹配并更新CSV文件
数据匹配拼接实现方案
核心实现思路:先预处理file2.txt构建编号到能力属性的映射关系(支持单个编号对应多个属性),再遍历file1.csv的每一行,用第4列的编号查询映射,将对应的所有属性追加到行末输出。
方案1:Bash(awk实现)
无额外依赖,处理大文件速度快,单命令即可完成:
awk ' # 先处理file2.txt,构建编号到属性的映射 NR==FNR { attrs[$1] = attrs[$1] " " $2 next } # 处理file1.csv,拼接属性输出 { print $0 (attrs[$4] ? attrs[$4] : "") } ' file2.txt file1.csv
如果需要对同一个编号的属性按字母排序(匹配示例输出的顺序),可以调整为:
awk ' NR==FNR { if (!arr[$1]) arr[$1] = $2 else arr[$1] = arr[$1] " " $2 next } { if (arr[$4]) { n = split(arr[$4], a, " ") asort(a) res = "" for(i=1;i<=n;i++) res = res " " a[i] print $0 res } else { print $0 } } ' file2.txt file1.csv
运行后需要保存结果的话,在命令末尾加 > output.csv 即可。
方案2:Python实现
逻辑直观,方便后续扩展修改:
# 构建编号到属性的映射 attr_map = {} with open("file2.txt", "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue idx, attr = line.split(maxsplit=1) if idx not in attr_map: attr_map[idx] = [] attr_map[idx].append(attr) # 处理csv文件输出结果 with open("file1.csv", "r", encoding="utf-8") as f: for line in f: line = line.rstrip("\n") if not line: print(line) continue parts = line.split() idx = parts[3] if idx in attr_map: # 需要属性按字母排序的话,把attr_map[idx]换成sorted(attr_map[idx])即可 print(f"{line} {' '.join(attr_map[idx])}") else: print(line)
运行脚本后需要保存结果的话,执行 python3 script.py > output.csv 即可。
内容的提问来源于stack exchange,提问作者user17315589
相关产品推荐
相关产品推荐

