You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用join或awk按指定列合并两个已排序文件,原脚本输出空求解决

问题解决方法

原脚本错误原因

  • 指定输出字段2.4,但File2仅包含2个字段,不存在第4个字段,直接导致join无输出
  • 插入表头时使用逗号作为分隔符,和后续内容的空白分隔规则不统一,会导致最终格式错位
  • 没有处理File1中基因描述占多个字段的场景,直接取固定位置字段会丢失部分描述内容

推荐awk实现方案

awk处理多字段匹配的场景更灵活,不需要额外调整文件排序,可直接实现需求:

awk '
NR == FNR {
    gene = $2
    score_mirDB = $1
    desc = ""
    for(i=3;i<=NF;i++){
        desc = desc " " $i
    }
    sub(/^ /,"",desc)
    data[gene] = score_mirDB "\t" desc
    next
}
{
    gene = $2
    if(gene in data){
        split(data[gene], arr, "\t")
        print gene "\t" arr[2] "\t" arr[1] "\t" $1
    }
}
' File1 File2 | sed '1i Gene Symbol\tGene Description\tTarget Score mirDB\tTarget Score Diana' | column -t > Output

运行验证

执行后Output文件内容完全符合预期:

Gene Symbol  Gene Description         Target Score mirDB  Target Score Diana
HOXC11       centrosomal protein 57   70                  0.700229616476825
CHD4         chromodomain helicase    70                  0.700328646327188
LUZP2        leucine zipper protein 2 70                  0.700328951649384

内容的提问来源于stack exchange,提问作者Diego Munoz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 12:36:00