You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AWK处理FASTA文件:移除GGTT模式及之前所有内容

Solution to Remove All GGTT Patterns and Preceding/Intervening Content in FASTA Files

Got it, let's fix this FASTA processing task for you. The key requirement here is to strip out all occurrences of GGTT, plus everything before the first GGTT and between multiple GGTT instances, leaving only the content that comes after the final GGTT in each sequence.

Awk is perfect for this because it easily handles FASTA's header-sequence structure, even if sequences span multiple lines. Here's a script that does exactly what you need:

# Process FASTA headers and sequences
/^>/ {
    # If we have a stored sequence from the previous entry, process it first
    if (seq != "") {
        # Replace everything up to and including the LAST occurrence of GGTT with empty string
        sub(/.*GGTT/, "", seq)
        print header "\n" seq
    }
    # Store the new header and reset the sequence buffer
    header = $0
    seq = ""
    next
}
# Accumulate sequence lines (handles multi-line sequences too)
{ seq = seq $0 }
# Process the last sequence entry after the end of the file
END {
    sub(/.*GGTT/, "", seq)
    print header "\n" seq
}

How to Use This

  1. Save the script above to a file (e.g., process_fasta.awk)
  2. Run it against your input FASTA file:
    awk -f process_fasta.awk your_input.fasta
    

Or run it as a one-liner without saving a script:

awk '/^>/ {if(seq!=""){sub(/.*GGTT/,"",seq);print header"\n"seq};header=$0;seq="";next} {seq=seq$0} END{sub(/.*GGTT/,"",seq);print header"\n"seq}' your_input.fasta

Why This Works

  • The regex .*GGTT matches any characters (including none) up to and including the final GGTT in the sequence. Replacing this with an empty string leaves only what comes after the last GGTT.
  • It handles multi-line sequences seamlessly by accumulating all sequence lines into a single buffer before processing.
  • It properly separates header lines from sequence lines, ensuring each entry is processed correctly.

Test with Your Sample Input

If we run this against your sample FASTA:

HEADER1 AACTGGTTACGTGGTTCTCT
HEADER2 GGTTTCTC
HEADER3 CCAGGTTTCGAGGGGTTACGGGGTA

We get exactly your desired output:

HEADER1 CTCT
HEADER2 TCTC
HEADER3 ACGGGGTA

内容的提问来源于stack exchange,提问作者Apex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 09:47:41