Bash脚本更新CSV时标题字段含逗号的处理问题
解决Bash脚本处理含逗号的CSV字段截断问题
问题描述
ABC公司的task1.sh脚本在处理accounts.csv时,当title字段包含逗号(如Director, Second Career Services),会被误判为字段分隔符,导致title被截断,后续字段错位,输出不符合预期。
问题根源
Bash原生的read命令配合IFS=','无法识别CSV中用双引号包裹的含逗号字段,直接将逗号作为字段分隔符拆分,导致字段读取错误。
解决方案
修改脚本的CSV读取逻辑,正确解析带双引号的含逗号字段;同时确保输出时,含逗号的字段用双引号包裹,符合CSV规范。以下是修改后的完整脚本:
#!/bin/bash function format_name() { local first_name last_name # 将全名拆分为名和姓,转换为小写 read -r first_name last_name <<<"$(echo "$1" | tr '[:upper:]' '[:lower:]')" # 将名和姓的首字母转换为大写 first_name=$(tr '[:lower:]' '[:upper:]' <<<"${first_name:0:1}")${first_name:1} last_name=$(tr '[:lower:]' '[:upper:]' <<<"${last_name:0:1}")${last_name:1} if [[ $last_name == *-* ]]; then hyphen_pos=$(expr index "$last_name" "-") last_name=${last_name:0:hyphen_pos}$(tr '[:lower:]' '[:upper:]' <<<"${last_name:$hyphen_pos:1}")${last_name:$hyphen_pos+1} fi echo "$first_name $last_name" } function is_exception() { local word=$1 local exceptions=("Office" "for" "and" "the" "of" "new") # 例外词列表 for exception in "${exceptions[@]}"; do if [[ "${word,,}" == "${exception,,}" ]]; then return 0 # 该词是例外词(忽略大小写匹配) fi done return 1 # 该词不是例外词 } function format_title() { local title="$1" local formatted_title="" local word # 处理标题中的每个词 for word in $title; do if is_exception "$word"; then # 例外词转换为小写 formatted_title+=" $(tr '[:upper:]' '[:lower:]' <<<"$word")" else # 非例外词首字母大写,其余小写 local lower_word=$(tr '[:upper:]' '[:lower:]' <<<"$word") formatted_title+=" $(tr '[:lower:]' '[:upper:]' <<<"${lower_word:0:1}")${lower_word:1}" fi done # 移除开头空格,若结果含逗号则包裹双引号 local result="${formatted_title:1}" if [[ "$result" == *","* ]]; then echo "\"$result\"" else echo "$result" fi } function generate_email() { local alias=$1 local location_id=$2 local domain="@abc.com" email=${alias} # 若别名重复,添加location_id if [[ ${alias_counts[$alias]} -gt 1 ]]; then email=${email}${location_id} fi email=${email}${domain} echo "$email" } function create_email_alias() { local first_name="${1%% *}" local last_name="${1#* }" alias=$(echo "${first_name:0:1}${last_name}" | tr '[:upper:]' '[:lower:]') echo "$alias" } if [ $# -ne 1 ]; then echo "Usage: $0 <accounts.csv>" exit 1 fi declare -A alias_counts alias_array=() input_file=$1 output_file="accounts_new.csv" # 使用awk解析CSV,处理带引号的字段,将每行转换为用|分隔的临时格式(假设数据中不含|) awk ' BEGIN { FPAT = "([^,]*)|(\"[^\"]+\")" OFS = "|" } { for (i=1; i<=NF; i++) { # 去掉字段的双引号 gsub(/^"|"$/, "", $i) } print $0 } ' "$input_file" > temp_processed.csv read -r header < temp_processed.csv # 把临时分隔符|改回逗号,写入输出文件 echo "${header//|/,}" >"$output_file" skip_header=true while IFS='|' read -r id location_id name title email department; do if $skip_header; then skip_header=false continue fi alias=$(create_email_alias "$name") alias_array+=("$alias") ((alias_counts[$alias]++)) done < temp_processed.csv skip_header=true while IFS='|' read -r id location_id name title email department; do if $skip_header; then skip_header=false continue fi # 格式化姓名和标题 formatted_name=$(format_name "$name") formatted_title=$(format_title "$title") formatted_email=$(generate_email "$(create_email_alias "$name")" "$location_id") # 将格式化后的数据写入输出文件,注意字段顺序 echo "$id,$location_id,$formatted_name,$formatted_title,$formatted_email,$department" >>"$output_file" done < temp_processed.csv # 清理临时文件 rm temp_processed.csv echo "Processing done. $output_file created."
关键修改说明
- CSV解析逻辑:使用
awk的FPAT特性正确识别带引号的含逗号字段,将每行转换为用|(假设数据中不含该字符)分隔的临时格式,避免逗号干扰字段读取。 - 标题格式化输出:在
format_title函数末尾添加判断,若处理后的标题含逗号,则用双引号包裹,符合CSV规范。 - 例外词匹配优化:修改
is_exception函数的匹配逻辑,忽略大小写,避免因输入大小写不一致导致匹配失败。 - 临时文件清理:处理完成后删除临时文件,避免残留。
测试验证
用提供的accounts.csv测试修改后的脚本,输出的accounts_new.csv中,含逗号的title字段会被正确包裹双引号,且内容完整,符合预期输出。
内容的提问来源于stack exchange,提问作者ASK
相关产品推荐
相关产品推荐

