在Bash中批量提取多行相同文件的指定制表符分隔列并合并
批量提取多文件指定列并按行合并
我找到不少处理2-3个文件的方案,但要处理30个文件时没找到合适方法,自己写循环又卡壳了,应该有更高效的办法。下面用3个测试文件验证需求,所有文件行数一致,目标是提取每个文件的第3列并按行合并。
测试文件内容
test1.txt
1 A D 2 B E 3 C F
test2.txt
1 G J 2 H K 3 I L
test3.txt
1 M P 2 N R 3 O S
期望输出(out.txt)
D J P E K R F L S
我尝试过的无效/有问题的命令
- 循环卡顿:创建空
out.txt后遍历文件,但逻辑错误导致卡壳$cat out.txt $for file in test* $do $cat > temp.txt $paste temp.txt <(cut -f3 $file) >> out.txt $done - 使用
test{2..3}.txt配合paste,test3的列被放到了第4-6行:$paste test1.txt <(cut -f3 test{2..3}.txt) >> out.txt - 合并所有文件成功,但无法只选特定列:
$paste -d'\t' test* >> out.txt - 生成额外行的无效命令:
$paste -d'\t' empty_file.txt <(cut -f3 test*) >> out.txt
可行解决方案
方法1:用awk一次性处理所有文件
awk可以同时读取多个文件,按行同步处理,是最高效的方案:
awk '{ # 把每个文件的第3列存入对应数组 col[FILENAME][NR] = $3 } END { # 遍历所有行 for (i=1; i<=NR; i++) { # 遍历所有文件,输出对应行的第3列 for (file in col) { printf "%s\t", col[file][i] } printf "\n" } }' test*.txt > out.txt
如果需要控制输出的文件顺序,可以明确指定文件名列表(比如test1.txt test2.txt test3.txt)代替test*.txt,避免通配符排序问题。
方法2:改进paste+cut的用法
直接给paste传递每个文件的第3列处理结果,不需要中间文件:
paste <(cut -f3 test1.txt) <(cut -f3 test2.txt) <(cut -f3 test3.txt) > out.txt
如果文件是按规律命名的(比如test1到test30),可以用循环生成所有<(cut...)段:
# 生成所有cut命令的参数 args=() for file in test*.txt; do args+=("<(cut -f3 $file)") done # 执行paste paste "${args[@]}" > out.txt
方法3:修复你的循环逻辑
原来的循环错误在于每次都从空的temp.txt读取,正确的逻辑应该是每次把out.txt作为输入,和新列合并后覆盖out.txt:
# 初始化out.txt为第一个文件的第3列 cut -f3 test1.txt > out.txt # 遍历剩下的文件 for file in test{2..3}.txt; do # 把当前out.txt和新文件的第3列合并,临时保存到temp.txt,再替换out.txt paste out.txt <(cut -f3 $file) > temp.txt mv temp.txt out.txt done
内容的提问来源于stack exchange,提问作者Eva Colla'kova'
相关产品推荐
相关产品推荐

