修复Bash脚本处理CSV时丢失department列数据的问题
修复Bash CSV处理脚本丢失department列的问题
问题核心
原脚本依赖简单逗号分割处理CSV,但部分行的department字段后无尾随逗号,导致awk解析时字段索引错位,最终丢失department数据。需保留姓名规范化、重复邮箱追加location_id的核心功能,同时修复字段解析问题。
原素材参考
原CSV内容
id,name,department,location_id 1,john doe,sales,1 2,jane smith,engineering 3,bob jones,marketing,2 4,john doe,finance,1
原脚本
#!/bin/bash awk -F ',' ' NR==1 {print $1","$2","$3","$4",""email"; next} { # 规范化姓名:首字母大写 split($2, name_arr, " "); formatted_name = toupper(substr(name_arr[1],1,1)) substr(name_arr[1],2) " " toupper(substr(name_arr[2],1,1)) substr(name_arr[2],2); # 生成邮箱 email = tolower(name_arr[1]) "." tolower(name_arr[2]) "@abc.com"; # 检查重复邮箱 if (email in seen) { email = email "." $4; } else { seen[email] = 1; } # 输出行 print $1","formatted_name","$3","$4","email; } ' input.csv > output.csv
当前错误输出
id,name,department,location_id,email 1,John Doe,sales,1,john.doe@abc.com 2,Jane Smith,,2,jane.smith@abc.com 3,Bob Jones,marketing,2,bob.jones@abc.com 4,John Doe,finance,1,john.doe@abc.com.1
期望输出
id,name,department,location_id,email 1,John Doe,sales,1,john.doe@abc.com 2,Jane Smith,engineering,,jane.smith@abc.com 3,Bob Jones,marketing,2,bob.jones@abc.com 4,John Doe,finance,1,john.doe@abc.com.1
修复后的脚本
#!/bin/bash awk -F ',' ' BEGIN {OFS=","} NR==1 {print $0",email"; next} { # 补全字段:确保每行有4个字段,缺失的用空字符串填充 for(i=NF+1; i<=4; i++) $i = ""; # 规范化姓名格式 split($2, name_arr, " "); formatted_name = toupper(substr(name_arr[1],1,1)) substr(name_arr[1],2) " " toupper(substr(name_arr[2],1,1)) substr(name_arr[2],2); # 生成基础邮箱 email = tolower(name_arr[1]) "." tolower(name_arr[2]) "@abc.com"; # 处理重复邮箱:仅当location_id非空时追加 if (email in seen) { if ($4 != "") email = email "." $4; } else { seen[email] = 1; } # 输出完整行 print $1, formatted_name, $3, $4, email; } ' input.csv > output.csv
关键修改说明
- 字段补全逻辑:通过循环将每行字段补全至原CSV的列数(4列),避免因行尾无逗号导致的
$3(department字段)索引错位。 - 统一输出分隔符:用
OFS=","替代手动拼接逗号,保证CSV格式规范。 - 空location_id处理:当location_id为空时,重复邮箱不再追加空字符串,避免生成无效邮箱格式。
内容的提问来源于stack exchange,提问作者Narek Arakelyan
相关产品推荐
相关产品推荐

