You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

UNIX脚本处理数据集标称属性独热编码输出异常如何解决

问题根因
  • 替换顺序错误:先处理了f、m这类单字符取值,导致包含这些字符的长属性值(如female、male、asympt)被提前截断替换,出现输出中的1 0ale、asy0 1p这类乱码
  • 匹配无边界:sed默认匹配子串,normal同时出现在restecg和thal两个属性中,f同时出现在female和fbs属性中,无边界替换会导致跨属性错误匹配
  • 拼写错误:cp属性取值non_anginal被误写为non_aginal,无法匹配对应值
  • 重复替换冲突:先后两次对f做全局替换,分别对应sex和fbs属性的取值,逻辑完全冲突
  • 缺失替换规则:num属性共5个取值,仅写了<50和>50_1的替换规则,剩下3个取值没有处理
修正方案

推荐用awk按字段处理,每列对应固定属性,完全避免跨属性匹配问题,修正后的脚本如下:

# 预处理原始数据
tail -n +19 ~file_location | fgrep -v "%" | shuf | sed -e "s/,/ /g" > temp1.txt

# 生成训练集文件
/bin/echo "SNNS pattern definition file V3.2"  > heart-v-train.pat
/bin/echo "generated at Mon Apr 25 15:58:23 1994"  >> heart-v-train.pat
/bin/echo ""  >> heart-v-train.pat
/bin/echo ""  >> heart-v-train.pat
/bin/echo "No. of patterns : 364"  >> heart-v-train.pat
/bin/echo "No. of input units : 25"  >> heart-v-train.pat
/bin/echo "No. of output units : 2"  >> heart-v-train.pat

# 按字段做独热编码
head -$TRAIN temp1.txt | awk '
# 定义各标称属性的编码映射
BEGIN {
    # sex字段编码
    sex_map["female"] = "1 0"; sex_map["male"] = "0 1"
    # cp字段编码
    cp_map["typ_angina"] = "1 0 0 0"; cp_map["asympt"] = "0 1 0 0"; cp_map["non_anginal"] = "0 0 1 0"; cp_map["atyp_angina"] = "0 0 0 1"
    # fbs字段编码
    fbs_map["t"] = "1 0"; fbs_map["f"] = "0 1"
    # restecg字段编码
    restecg_map["left_vent_hyper"] = "1 0 0"; restecg_map["normal"] = "0 1 0"; restecg_map["st_t_wave_abnormality"] = "0 0 1"
    # exang字段编码
    exang_map["no"] = "1 0"; exang_map["yes"] = "0 1"
    # slope字段编码
    slope_map["up"] = "1 0 0"; slope_map["flat"] = "0 1 0"; slope_map["down"] = "0 0 1"
    # thal字段编码
    thal_map["fixed_defect"] = "1 0 0"; thal_map["normal"] = "0 1 0"; thal_map["reversable_defect"] = "0 0 1"
    # num字段编码
    num_map["<50"] = "1 0 0 0 0"; num_map[">50_1"] = "0 1 0 0 0"; num_map[">50_2"] = "0 0 1 0 0"; num_map[">50_3"] = "0 0 0 1 0"; num_map[">50_4"] = "0 0 0 0 1"
}
{
    # 按字段输出,数值字段直接输出,标称字段替换为编码
    print $1, sex_map[$2], cp_map[$3], $4, $5, fbs_map[$6], restecg_map[$7], $8, exang_map[$9], $10, slope_map[$11], $12, thal_map[$13], num_map[$14]
}' >> heart-v-train.pat

如果要继续用sed方案,需要调整替换顺序为长值优先,并添加单词边界\b避免子串匹配,修正拼写和重复替换问题:

head -$TRAIN temp1.txt |
    # 先替换最长的属性值
    sed -e "s/\bleft_vent_hyper\b/1 0 0/g" \
        -e "s/\bst_t_wave_abnormality\b/0 0 1/g" \
        -e "s/\breversable_defect\b/0 0 1/g" \
        -e "s/\bfixed_defect\b/1 0 0/g" \
        -e "s/\btyp_angina\b/1 0 0 0/g" \
        -e "s/\basympt\b/0 1 0 0/g" \
        -e "s/\bnon_anginal\b/0 0 1 0/g" \
        -e "s/\batyp_angina\b/0 0 0 1/g" \
        -e "s/\bfemale\b/1 0/g" \
        -e "s/\bmale\b/0 1/g" \
        -e "s/\bnormal\b/0 1 0/g" \
        -e "s/\bno\b/1 0/g" \
        -e "s/\byes\b/0 1/g" \
        -e "s/\bup\b/1 0 0/g" \
        -e "s/\bflat\b/0 1 0/g" \
        -e "s/\bdown\b/0 0 1/g" \
        -e "s/\bt\b/1 0/g" \
        -e "s/\bf\b/0 1/g" \
        -e "s/\b<50\b/1 0 0 0 0/g" \
        -e "s/\b>50_1\b/0 1 0 0 0/g" \
        -e "s/\b>50_2\b/0 0 1 0 0/g" \
        -e "s/\b>50_3\b/0 0 0 1 0/g" \
        -e "s/\b>50_4\b/0 0 0 0 1/g" >> heart-v-train.pat

内容的提问来源于stack exchange,提问作者fRo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 01:48:05