BASH实现列表列规范化:将列表项拆分为独立行
问题:BASH中拆分TSV表格的逗号分隔列到多行
我想在BASH里把TSV表格中用逗号分隔的列表项拆成单独的行,比如把下面的TSV:
TYPE NAME Fruit apple,strawberry Vegetable potato
转换成:
TYPE NAME Fruit apple Fruit strawberry Vegetable potato
我写了下面的脚本,但运行后输出文件是空的,想知道问题在哪,有没有更优的处理方案?
#!/bin/bash # define the name of the input file input_file="plants.tsv" # define the name of the output file output_file="normalized_plants.tsv" # define the index of the list column (counting from 1) list_column=2 # create a new file with the headers for the output table head -n 1 "$input_file" > "$output_file" # read each line of the input file tail -n +2 "$input_file" | while IFS=$'\t' read -r line; do # extract the values for the list column list_values=$(echo "$line" | awk -F$'\t' '{print $'"$list_column"'}' | tr ',' '\n') # iterate over each value in the list column echo "$line" | awk -F$'\t' -v OFS=$'\t' -v list_column="$list_column" -v list_values="$list_values" ' NR == 1 { next } # skip the header row { split(list_values, values, "\n") for (i in values) { $list_column = values[i] print $0 } }' >> "$output_file" done
脚本问题分析
- NR判断逻辑错误:你通过
echo "$line"把单行传给awk,此时awk只处理这一行,NR == 1 { next }会直接跳过该行,导致完全没有输出。 - 换行符传递异常:
list_values中的换行符在传给awk时,可能被解析成无效分隔符,导致split无法正确拆分出目标值。 - 多进程冗余调用:每行都重复调用awk、tr等外部命令,不仅效率低,还容易引发变量传递的意外问题。
更优解决方案
方案1:纯awk处理(推荐)
awk天生适合文本处理,一次扫描即可完成所有操作,逻辑清晰且效率最高:
#!/usr/bin/awk -f BEGIN { FS = "\t" # 设置输入分隔符为制表符 OFS = "\t" # 设置输出分隔符为制表符 } NR == 1 { print $0 # 直接输出表头行 next } { split($2, name_list, ",") # 拆分第2列的逗号分隔值 for (i in name_list) { $2 = name_list[i] print $0 } }
使用步骤:
- 将上述代码保存为
split_tsv.awk - 赋予执行权限:
chmod +x split_tsv.awk - 运行脚本:
./split_tsv.awk plants.tsv > normalized_plants.tsv
如果需要动态指定拆分列(而非固定第2列),可以通过变量传递实现:
#!/usr/bin/awk -f BEGIN { FS = "\t" OFS = "\t" if (!target_col) target_col = 2 # 默认拆分第2列 } NR == 1 { print $0 next } { split($target_col, val_list, ",") for (i in val_list) { $target_col = val_list[i] print $0 } }
运行时指定列号:./split_tsv.awk -v target_col=2 plants.tsv > normalized_plants.tsv
方案2:简化版Bash脚本
如果更习惯用Bash处理,可简化逻辑,避免嵌套外部命令:
#!/bin/bash input_file="plants.tsv" output_file="normalized_plants.tsv" target_col=2 # 这里因为是TSV,前两列直接对应type和names,所以直接读取 # 先写入表头 head -n 1 "$input_file" > "$output_file" # 逐行处理数据行 tail -n +2 "$input_file" | while IFS=$'\t' read -r type names; do # 将逗号分隔的字符串拆分为数组 IFS=',' read -ra name_array <<< "$names" # 遍历数组输出每一行 for name in "${name_array[@]}"; do echo -e "$type\t$name" >> "$output_file" done done
这个脚本直接利用Bash内置的字符串拆分功能,无需频繁调用外部工具,逻辑更直观易懂。
内容的提问来源于stack exchange,提问作者sasa.asa
相关产品推荐
相关产品推荐

