You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWK中替换description字段换行符以修正CSV输出

问题:AWK生成CSV时因description字段含换行符导致格式异常

原脚本用于从结构化文本提取数据生成CSV,但当description字段包含换行时,输出出现多余换行,不符合预期格式:

#Removing single quote from file to avoid conflict with delmiter 
roledetails="${roledetails//","/"|"}" 
    awk '
    BEGIN                      { FS  = ":[[:space:]]+"
                                 OFS = ","
                               }
    $1 ~ /^name$/              { coll = $2 } #Collection Role Name
    $1 ~ /^description$/       { desc = $2} #Collection Role Desc
    $1 ~ / name$/              { name = $2 } #Single Role Name
    $1 ~ / roleTemplateAppId$/ { id = $2 } #Single Role App ID with all other field in last column
    $1 ~ / roleTemplateName$/  { print coll,desc,name,id,$2} #Single Role Template Name
    ' <<< "${roledetails}" >> "${fullfilepath}"

尝试用$1 ~ /^description$/ { desc = gsub(/\n/,"",$2)}解决无效,当前输出会把coll,desc单独占一行,后续role信息另起一行,而期望每一条role对应完整的一行CSV。


修正后的脚本
#Removing single quote from file to avoid conflict with delmiter 
roledetails="${roledetails//","/"|"}" 
    awk '
    BEGIN                      { FS  = ":[[:space:]]+"
                                 OFS = ","
                                 in_desc = 0
                               }
    # 处理顶层name字段
    $1 ~ /^name$/ { 
        coll = $2 
        next
    }
    # 处理顶层description,支持多行内容
    $1 ~ /^description$/ { 
        desc = substr($0, index($0, ":")+2) # 提取冒号后的全部内容(包括可能的换行后续行)
        in_desc = 1
        next
    }
    # 若处于description多行内容中,且当前行不是新的字段(不含冒号),则追加内容
    in_desc && !/:[[:space:]]+/ {
        desc = desc $0
        next
    }
    # 退出description多行模式,处理其他字段
    in_desc {
        gsub(/\n/, "", desc) # 移除所有换行符
        in_desc = 0
    }
    # 处理单个role的name字段
    $1 ~ / name$/ { 
        name = $2 
        next
    }
    # 处理单个role的AppID字段
    $1 ~ / roleTemplateAppId$/ { 
        id = $2 
        next
    }
    # 处理单个role的TemplateName,此时输出完整一行
    $1 ~ / roleTemplateName$/ { 
        print coll, desc, name, id, $2
        # 重置单个role的变量,避免残留
        name = id = ""
    }
    ' <<< "${roledetails}" >> "${fullfilepath}"

关键修改说明
  1. 多行description处理:
    原脚本默认按行读取,无法捕获跨多行的description内容。新增in_desc标记,当匹配到^description$时,提取该行冒号后的全部内容,并标记进入多行收集模式;后续若遇到不含冒号的行(属于description的换行内容),则追加到desc变量,直到遇到新的字段行才退出多行模式。

  2. 换行符移除时机:
    在退出多行模式时,统一对desc执行gsub(/\n/, "", desc),彻底移除所有换行符,确保desc是单行内容。

  3. 输出逻辑优化:
    仅当收集到roleTemplateName时才输出完整一行CSV,避免因description换行导致的提前输出,保证每一行都包含完整的coll, desc, name, id, roleTemplateName字段。


内容的提问来源于stack exchange,提问作者Lord OfTheRing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 21:55:25