You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BASH实现列表列规范化:将列表项拆分为独立行

问题:BASH中拆分TSV表格的逗号分隔列到多行

我想在BASH里把TSV表格中用逗号分隔的列表项拆成单独的行,比如把下面的TSV:

TYPE    NAME  
Fruit   apple,strawberry
Vegetable   potato

转换成:

TYPE    NAME  
Fruit   apple
Fruit   strawberry
Vegetable   potato

我写了下面的脚本,但运行后输出文件是空的,想知道问题在哪,有没有更优的处理方案?

#!/bin/bash

# define the name of the input file
input_file="plants.tsv"

# define the name of the output file
output_file="normalized_plants.tsv"

# define the index of the list column (counting from 1)
list_column=2

# create a new file with the headers for the output table
head -n 1 "$input_file" > "$output_file"

# read each line of the input file
tail -n +2 "$input_file" | while IFS=$'\t' read -r line; do
  # extract the values for the list column
  list_values=$(echo "$line" | awk -F$'\t' '{print $'"$list_column"'}' | tr ',' '\n')
  # iterate over each value in the list column
  echo "$line" | awk -F$'\t' -v OFS=$'\t' -v list_column="$list_column" -v list_values="$list_values" '
    NR == 1 { next } # skip the header row
    { 
      split(list_values, values, "\n")
      for (i in values) {
        $list_column = values[i]
        print $0
      }
    }' >> "$output_file"
done

脚本问题分析

  • NR判断逻辑错误:你通过echo "$line"把单行传给awk,此时awk只处理这一行,NR == 1 { next }会直接跳过该行,导致完全没有输出。
  • 换行符传递异常:list_values中的换行符在传给awk时,可能被解析成无效分隔符,导致split无法正确拆分出目标值。
  • 多进程冗余调用:每行都重复调用awk、tr等外部命令,不仅效率低,还容易引发变量传递的意外问题。

更优解决方案

方案1:纯awk处理(推荐)

awk天生适合文本处理,一次扫描即可完成所有操作,逻辑清晰且效率最高:

#!/usr/bin/awk -f
BEGIN {
    FS = "\t"  # 设置输入分隔符为制表符
    OFS = "\t" # 设置输出分隔符为制表符
}
NR == 1 {
    print $0  # 直接输出表头行
    next
}
{
    split($2, name_list, ",")  # 拆分第2列的逗号分隔值
    for (i in name_list) {
        $2 = name_list[i]
        print $0
    }
}

使用步骤:

  1. 将上述代码保存为split_tsv.awk
  2. 赋予执行权限:chmod +x split_tsv.awk
  3. 运行脚本:./split_tsv.awk plants.tsv > normalized_plants.tsv

如果需要动态指定拆分列(而非固定第2列),可以通过变量传递实现:

#!/usr/bin/awk -f
BEGIN {
    FS = "\t"
    OFS = "\t"
    if (!target_col) target_col = 2  # 默认拆分第2列
}
NR == 1 {
    print $0
    next
}
{
    split($target_col, val_list, ",")
    for (i in val_list) {
        $target_col = val_list[i]
        print $0
    }
}

运行时指定列号:./split_tsv.awk -v target_col=2 plants.tsv > normalized_plants.tsv

方案2:简化版Bash脚本

如果更习惯用Bash处理,可简化逻辑,避免嵌套外部命令:

#!/bin/bash

input_file="plants.tsv"
output_file="normalized_plants.tsv"
target_col=2  # 这里因为是TSV,前两列直接对应type和names,所以直接读取

# 先写入表头
head -n 1 "$input_file" > "$output_file"

# 逐行处理数据行
tail -n +2 "$input_file" | while IFS=$'\t' read -r type names; do
    # 将逗号分隔的字符串拆分为数组
    IFS=',' read -ra name_array <<< "$names"
    # 遍历数组输出每一行
    for name in "${name_array[@]}"; do
        echo -e "$type\t$name" >> "$output_file"
    done
done

这个脚本直接利用Bash内置的字符串拆分功能,无需频繁调用外部工具,逻辑更直观易懂。


内容的提问来源于stack exchange,提问作者sasa.asa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 22:23:28