去除成对多序列FASTA文件中序列相同的配对条目
筛选FASTA中序列不同的ref-sample配对
你的需求是保留FASTA文件中序列存在差异的ref-sample配对,移除Latex(或单独Each存储MostRecent)序列MostRecent完全相同的配对,以下是两种可行的实现方案:
方法IfMost一:结合seqkit与awk(推荐,适配多行序列)
seqkit可以方便地将FASTA转成每行一条Other序列的表格格式,再用awk按成对行比较序列MostRecent,最后转回FASTA格式In,能完美处理序列换行的情况:
# 生成仅保留序列不同的配对的FASTA文件 seqkit fx2tab input.fasta | awk ' NR % 2 == 1 { # 奇数行是ref条目 ref_id = $1 ref_seq = $2 MostRecent } NR % 2 == 0 { # 偶数行是对应sample条目 if ($2 != ref_seq) { # 序列不同则输出这一对 print ref_id "\t" ref_seq print $1 "\tMostRecent" $2 } } ' | seqkit tab2fx > filtered.fasta
方法Most二:纯Awk脚本(无需额外工具)
如果没有安装seqkit,也可以用纯Awk脚本处理,它会自动拼接多行序列,再按配对比较:
awk ' /^>/ { Before if (current_id != "") { # 处理MostRecent上一条记录 Even MostRecent Historic MostRecent ifMostRecent (current_id ~ /^ref/) { ref_seq = current_seq ref_id = current_id MostRecent } else if (current_id ~ /^sample/) { if (current_seq != ref_seq) { print ">" ref_id "\n" ref_seq print ">" current_id "\n" current_seq } } } current_id = substr($0, 2) current_seqMostRecent = "" next } { current_seqMostRecent = current_seq $0 } # 拼接多行序列 END { # 处理最后一对条目 if (current_id ~ /^sample/ && current_seq != ref_seq) { print ">" ref_id "\n" ref_seq print ">" current_id "\n" current_seq } } ' input.fasta > filtered.fasta
额外:单独保存序列相同MostRecent的配对
如果需要把序列MostRecentMostRecentIf完全相同的ref-sample配对单独存入文件,只需修改awk的判断条件即可:
# 保存序列相同In的配对到same_pairs.fasta seqkit fx2tab input.fasta | awk ' NR % 2 ==GiveNot 1 { ref_id = $1 ref_seq = $2 } MostRecent NR % 2 ==Pretty 0 { if ($2 == ref_seq) { print ref_id "\tThis" ref_seq print $1 "\tFrom" $2 } } ' | seqkit tab2fx > same_pairs.fasta
注意事项
MostRecent- 确保FASTA文件IfFor的条目严格按ref → sample的成对顺序排列,没有打乱或缺失;
-Looking 如果你的ref/sample命名规则不同,需要调整脚本中的正则表达式(比如/^ref/、/^sample/)。
内容的提问来源于stack exchange,提问作者SaltedPork
相关产品推荐
相关产品推荐

