You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

根据TSV文件替换FASTA头部匹配ARO字段的对应内容

Hey there! As a microbiology researcher with basic Bash/Python skills, I’ve put together a simple, easy-to-follow Python script that handles exactly what you need: swapping out the ARO field in your FASTA headers with the matching gene name and functional description from your TSV file when there’s a match. No fancy bioinformatics tools required—just plain Python you can tweak if needed.

How it works

First, we'll load your TSV data into a quick-lookup dictionary, then iterate through your FASTA file to modify headers where an ARO match exists. Here's the full script:

def main():
    import sys
    # Make sure we have all required input files and output path
    if len(sys.argv) != 4:
        print("Usage: python replace_fasta_headers.py <input_fasta> <input_tsv> <output_fasta>")
        sys.exit(1)
    
    fasta_in = sys.argv[1]
    tsv_in = sys.argv[2]
    fasta_out = sys.argv[3]
    
    # Step 1: Build a map of ARO IDs to their gene name + description
    aro_lookup = {}
    with open(tsv_in, 'r', encoding='utf-8') as tsv_file:
        for line in tsv_file:
            # Split TSV line into its three components
            line_parts = line.strip().split('\t')
            if len(line_parts) == 3:
                aro_id, gene_name, description = line_parts
                aro_lookup[aro_id] = f"{gene_name} {description}"
    
    # Step 2: Process the FASTA file line by line
    with open(fasta_in, 'r', encoding='utf-8') as fasta_file, open(fasta_out, 'w', encoding='utf-8') as out_file:
        for line in fasta_file:
            stripped_line = line.strip()
            if stripped_line.startswith('>'):
                # Handle header lines
                header = stripped_line[1:]  # Remove the leading ">"
                header_segments = header.split('|')
                new_segments = []
                
                for segment in header_segments:
                    if segment.startswith('ARO:'):
                        # Check if this ARO has a match in our lookup map
                        if segment in aro_lookup:
                            new_segments.append(aro_lookup[segment])
                        else:
                            # No match found, keep original ARO segment
                            new_segments.append(segment)
                    else:
                        # Keep non-ARO segments exactly as they are
                        new_segments.append(segment)
                
                # Reconstruct the modified header and write it
                new_header = '>' + '|'.join(new_segments)
                out_file.write(new_header + '\n')
            else:
                # Write sequence lines without any changes
                out_file.write(stripped_line + '\n')
    
    print(f"Success! Modified FASTA saved to {fasta_out}")

if __name__ == "__main__":
    main()

Step-by-step breakdown

  1. Argument Check: The script first makes sure you've provided all three required files (input FASTA, input TSV, output FASTA) so you don't run it incorrectly.
  2. TSV Lookup Dictionary: We read your TSV file and store each ARO ID as a key, with the combined gene name + description as its value. This lets us quickly check for matches without scanning the entire TSV every time.
  3. FASTA Processing: We loop through every line of your FASTA file:
    • For header lines (starting with >), we split the header into segments separated by |.
    • If a segment starts with ARO:, we check if it exists in our lookup dictionary. If it does, we replace the ARO segment with the gene name + description; if not, we leave it as-is.
    • Sequence lines are written exactly as they are—no changes to your actual genetic data.

How to run it

  1. Save the script as replace_fasta_headers.py
  2. Open your terminal and run:
    python replace_fasta_headers.py your_input.fasta your_aro_data.tsv your_output.fasta
    

Example output

Using your sample files, the script will turn this header:

Prevalence_Sequence_ID:1|ARO:3003072|RES:mphL|Protein Homolog Model

Into this:

Prevalence_Sequence_ID:1|mphL mphL是一种染色体编码的大环内酯磷酸转移酶,可灭活14元和15元大环内酯类药物,如红霉素、克拉霉素、阿奇霉素。|RES:mphL|Protein Homolog Model

All unmatched headers will stay exactly the same.

内容的提问来源于stack exchange,提问作者CptChiler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:15:23