如何调整Bash中SED/AWK/Perl输出,合并行并生成指定CSV格式?
解决方案
一、先处理客户条目格式(合并多行+去除多余空格)
用awk处理嵌套结构文本比sed更适合多行合并场景,以下命令可直接输出格式化后的客户条目:
awk '/customers {/ {in_cust=1; next} in_cust && /}/ {in_cust=0; next} in_cust { if ($0 ~ /{/) { name = $1 brace_part = gensub(/[[:space:]]*({.*)/, "\\1", "g", substr($0, index($0, "{"))) curr = name brace_part gsub(/[[:space:]]+/, "", curr) in_entry=1 } else if ($0 ~ /}/) { curr = curr gensub(/[[:space:]]+/, "", "g", $0) print curr in_entry=0 } else if (in_entry) { line = gensub(/^[[:space:]]+/, "", "g", $0) curr = curr line } }' input.txt
输出结果:
mary{ } freddy{ } bob{spouse betty}
二、生成目标CSV文件
扩展上述awk脚本,同时提取产品名称并整合所有字段生成CSV:
BEGIN { print "product,customers,another_column" } /name {/ { product = gensub(/.*name {[[:space:]]*([^}]+)[[:space:]]*}.*/, "\\1", "g", $0) } /customers {/ {in_cust=1; next} in_cust && /}/ { in_cust=0 gsub(/^[[:space:]]+/, "", cust_list) next } in_cust { if ($0 ~ /{/) { name = $1 brace_part = gensub(/[[:space:]]*({.*)/, "\\1", "g", substr($0, index($0, "{"))) curr_cust = name brace_part gsub(/[[:space:]]+/, "", curr_cust) in_entry=1 } else if ($0 ~ /}/) { curr_cust = curr_cust gensub(/[[:space:]]+/, "", "g", $0) cust_list = cust_list (cust_list ? " " : "") curr_cust in_entry=0 } else if (in_entry) { line = gensub(/^[[:space:]]+/, "", "g", $0) curr_cust = curr_cust line } } END { print product "," cust_list ",something_else" } ' input.txt > output.csv
运行后output.csv内容:
product,customers,another_column thing1,mary{ } freddy{ } bob{spouse betty},something_else
三、关键逻辑说明
- 块状态跟踪:用
in_cust和in_entry标记变量,分别判断是否进入customers块、是否处于单个客户的条目块。 - 产品名提取:匹配
name { ... }行,通过正则提取{}内的产品名称。 - 客户条目处理:
- 客户开头行:提取名称和
{,去除中间空格。 - 客户内容行:去掉开头缩进后追加到当前条目。
- 客户结束行:去掉空格后追加,完成条目并加入客户列表。
- 客户开头行:提取名称和
- CSV生成:开头打印表头,结尾拼接所有字段输出CSV行。
四、扩展提示
如果文本包含多个product块,或another_column需要从文本提取,可调整脚本:
- 在每个
product块结束时输出一行CSV,而非仅在END阶段输出。 - 参考产品名提取逻辑,增加
another_column的匹配和提取代码。
内容的提问来源于stack exchange,提问作者hubdows
相关产品推荐
相关产品推荐

