You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sed、awk或grep提取字符串中的<<符号及对应字母

提取DNA协方差模型中<<区域对应的序列与符号

问题说明

现有一份DNA协方差模型数据文件,每行格式为ID : 序列/协方差符号,输入示例如下:

NC_013791.2.2 : GCTCAGCTGGCtAGAG
NC_013791.2.2 : >>>>.........<<<
NC_013791.2.3 : GCTCAGCTGGCtAGAG
NC_013791.2.3 : >>>>..<<<<......
NC_013791.2.4 : GCTCAGCTGGCtAGGA
NC_013791.2.4 : >>>>.........<<<
NC_013791.2.5 : GCTCAGCTGACtACAG
NC_013791.2.5 : >>>>..<<<<......

需要提取每个ID对应的<<符号区域的序列及符号本身,期望输出:

NC_013791.2.2 :  GAG
NC_013791.2.2 :  <<<
NC_013791.2.3 : CTGG
NC_013791.2.3 : <<<<
NC_013791.2.4 : GGA
NC_013791.2.4 : <<<
NC_013791.2.5 : CTGA
NC_013791.2.5 : <<<<

解决方案

1. awk 方案(推荐,逻辑清晰)

通过存储每个ID的序列,遇到协方差行时定位<<区域的位置和长度,再从对应序列中截取片段:

awk '
    /: [ACGTacgt]/ {
        id = $1 " : "
        seq = substr($0, index($0, ": ") + 2)
        seq_map[id] = seq
        next
    }
    /: [><.]/ {
        id = $1 " : "
        sym = substr($0, index($0, ": ") + 2)
        start = index(sym, "<<")
        len = 0
        while (substr(sym, start + len, 1) == "<") len++
        target_seq = substr(seq_map[id], start, len)
        printf "%s %s\n", id, target_seq
        printf "%s %s\n", id, substr(sym, start, len)
    }
' input.txt

2. sed 方案(纯文本处理,依赖行顺序)

利用sed的暂存功能,成对处理序列行和协方差行,定位<<区域后截取对应内容:

sed -n '
    /: [ACGTacgt]/{
        h
        s/^.*: //p
        g
        s/: .*//
        s/$/ : /
        H
        d
    }
    /: [><.]/{
        s/^.*: //
        h
        s/.*<<//
        s/[^<].*//
        len=length($0)
        g
        s/<<.*//
        start=length($0)+1
        x
        G
        s/\(.*\)\n\(.*\)/echo "\2" | cut -c'"$start-$((start+len-1))"'/e
        x
        s/.*\(<<\{1,\}\).*/\1/
        G
        s/\(.*\)\n\(.*: \)/\2\1/
        p
    }
' input.txt

3. Bash 脚本方案(结合grep/cut,易理解)

逐行成对读取序列和协方差行,通过shell命令提取目标区域:

#!/bin/bash
while read -r seq_line && read -r sym_line; do
    id=$(echo "$seq_line" | cut -d: -f1)
    seq=$(echo "$seq_line" | cut -d: -f2 | xargs)
    sym=$(echo "$sym_line" | cut -d: -f2 | xargs)
    
    start=$(expr index "$sym" '<<')
    len=$(echo "$sym" | sed 's/.*<<//; s/[^<].*//' | wc -c)
    len=$((len-1))
    
    target_seq=$(echo "$seq" | cut -c"$start-$((start+len-1))")
    target_sym=$(echo "$sym" | cut -c"$start-$((start+len-1))")
    
    echo "$id :  $target_seq"
    echo "$id :  $target_sym"
done < input.txt

内容的提问来源于stack exchange,提问作者Pandu C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 03:58:12