如何用sed、awk或grep提取字符串中的<<符号及对应字母
提取DNA协方差模型中<<区域对应的序列与符号
问题说明
现有一份DNA协方差模型数据文件,每行格式为ID : 序列/协方差符号,输入示例如下:
NC_013791.2.2 : GCTCAGCTGGCtAGAG NC_013791.2.2 : >>>>.........<<< NC_013791.2.3 : GCTCAGCTGGCtAGAG NC_013791.2.3 : >>>>..<<<<...... NC_013791.2.4 : GCTCAGCTGGCtAGGA NC_013791.2.4 : >>>>.........<<< NC_013791.2.5 : GCTCAGCTGACtACAG NC_013791.2.5 : >>>>..<<<<......
需要提取每个ID对应的<<符号区域的序列及符号本身,期望输出:
NC_013791.2.2 : GAG NC_013791.2.2 : <<< NC_013791.2.3 : CTGG NC_013791.2.3 : <<<< NC_013791.2.4 : GGA NC_013791.2.4 : <<< NC_013791.2.5 : CTGA NC_013791.2.5 : <<<<
解决方案
1. awk 方案(推荐,逻辑清晰)
通过存储每个ID的序列,遇到协方差行时定位<<区域的位置和长度,再从对应序列中截取片段:
awk ' /: [ACGTacgt]/ { id = $1 " : " seq = substr($0, index($0, ": ") + 2) seq_map[id] = seq next } /: [><.]/ { id = $1 " : " sym = substr($0, index($0, ": ") + 2) start = index(sym, "<<") len = 0 while (substr(sym, start + len, 1) == "<") len++ target_seq = substr(seq_map[id], start, len) printf "%s %s\n", id, target_seq printf "%s %s\n", id, substr(sym, start, len) } ' input.txt
2. sed 方案(纯文本处理,依赖行顺序)
利用sed的暂存功能,成对处理序列行和协方差行,定位<<区域后截取对应内容:
sed -n ' /: [ACGTacgt]/{ h s/^.*: //p g s/: .*// s/$/ : / H d } /: [><.]/{ s/^.*: // h s/.*<<// s/[^<].*// len=length($0) g s/<<.*// start=length($0)+1 x G s/\(.*\)\n\(.*\)/echo "\2" | cut -c'"$start-$((start+len-1))"'/e x s/.*\(<<\{1,\}\).*/\1/ G s/\(.*\)\n\(.*: \)/\2\1/ p } ' input.txt
3. Bash 脚本方案(结合grep/cut,易理解)
逐行成对读取序列和协方差行,通过shell命令提取目标区域:
#!/bin/bash while read -r seq_line && read -r sym_line; do id=$(echo "$seq_line" | cut -d: -f1) seq=$(echo "$seq_line" | cut -d: -f2 | xargs) sym=$(echo "$sym_line" | cut -d: -f2 | xargs) start=$(expr index "$sym" '<<') len=$(echo "$sym" | sed 's/.*<<//; s/[^<].*//' | wc -c) len=$((len-1)) target_seq=$(echo "$seq" | cut -c"$start-$((start+len-1))") target_sym=$(echo "$sym" | cut -c"$start-$((start+len-1))") echo "$id : $target_seq" echo "$id : $target_sym" done < input.txt
内容的提问来源于stack exchange,提问作者Pandu C
相关产品推荐
相关产品推荐

