根据TSV文件替换FASTA头部匹配ARO字段的对应内容
Hey there! As a microbiology researcher with basic Bash/Python skills, I’ve put together a simple, easy-to-follow Python script that handles exactly what you need: swapping out the ARO field in your FASTA headers with the matching gene name and functional description from your TSV file when there’s a match. No fancy bioinformatics tools required—just plain Python you can tweak if needed.
How it works
First, we'll load your TSV data into a quick-lookup dictionary, then iterate through your FASTA file to modify headers where an ARO match exists. Here's the full script:
def main(): import sys # Make sure we have all required input files and output path if len(sys.argv) != 4: print("Usage: python replace_fasta_headers.py <input_fasta> <input_tsv> <output_fasta>") sys.exit(1) fasta_in = sys.argv[1] tsv_in = sys.argv[2] fasta_out = sys.argv[3] # Step 1: Build a map of ARO IDs to their gene name + description aro_lookup = {} with open(tsv_in, 'r', encoding='utf-8') as tsv_file: for line in tsv_file: # Split TSV line into its three components line_parts = line.strip().split('\t') if len(line_parts) == 3: aro_id, gene_name, description = line_parts aro_lookup[aro_id] = f"{gene_name} {description}" # Step 2: Process the FASTA file line by line with open(fasta_in, 'r', encoding='utf-8') as fasta_file, open(fasta_out, 'w', encoding='utf-8') as out_file: for line in fasta_file: stripped_line = line.strip() if stripped_line.startswith('>'): # Handle header lines header = stripped_line[1:] # Remove the leading ">" header_segments = header.split('|') new_segments = [] for segment in header_segments: if segment.startswith('ARO:'): # Check if this ARO has a match in our lookup map if segment in aro_lookup: new_segments.append(aro_lookup[segment]) else: # No match found, keep original ARO segment new_segments.append(segment) else: # Keep non-ARO segments exactly as they are new_segments.append(segment) # Reconstruct the modified header and write it new_header = '>' + '|'.join(new_segments) out_file.write(new_header + '\n') else: # Write sequence lines without any changes out_file.write(stripped_line + '\n') print(f"Success! Modified FASTA saved to {fasta_out}") if __name__ == "__main__": main()
Step-by-step breakdown
- Argument Check: The script first makes sure you've provided all three required files (input FASTA, input TSV, output FASTA) so you don't run it incorrectly.
- TSV Lookup Dictionary: We read your TSV file and store each ARO ID as a key, with the combined gene name + description as its value. This lets us quickly check for matches without scanning the entire TSV every time.
- FASTA Processing: We loop through every line of your FASTA file:
- For header lines (starting with
>), we split the header into segments separated by|. - If a segment starts with
ARO:, we check if it exists in our lookup dictionary. If it does, we replace the ARO segment with the gene name + description; if not, we leave it as-is. - Sequence lines are written exactly as they are—no changes to your actual genetic data.
- For header lines (starting with
How to run it
- Save the script as
replace_fasta_headers.py - Open your terminal and run:
python replace_fasta_headers.py your_input.fasta your_aro_data.tsv your_output.fasta
Example output
Using your sample files, the script will turn this header:
Prevalence_Sequence_ID:1|ARO:3003072|RES:mphL|Protein Homolog Model
Into this:
Prevalence_Sequence_ID:1|mphL mphL是一种染色体编码的大环内酯磷酸转移酶,可灭活14元和15元大环内酯类药物,如红霉素、克拉霉素、阿奇霉素。|RES:mphL|Protein Homolog Model
All unmatched headers will stay exactly the same.
内容的提问来源于stack exchange,提问作者CptChiler

