基于UniprotID及目标残基批量获取±6位13mer肽段的方法咨询
Hey Emanuele, great question! Here are a few practical ways to batch extract those 13mer peptides (target residue plus 6 amino acids before and after) from your list of Uniprot IDs and target positions. I’ll break down options for Python, R, and no-code online tools to fit your workflow:
The easiest way is to use the Biopython library, which lets you fetch Uniprot sequences directly and manipulate them. Here's a step-by-step example:
First, install Biopython if you haven’t already:
pip install biopythonThen use this script to process your list:
from Bio import SwissProt def extract_13mer(uniprot_id, target_residue, target_pos): # Handle isoform IDs by stripping the suffix if needed clean_id = uniprot_id.split('_')[0] handle = SwissProt.get_sprot_raw(clean_id) record = SwissProt.read(handle) seq = record.sequence # Uniprot positions are 1-based; Python strings are 0-based start = max(0, target_pos - 7) end = min(len(seq), target_pos + 6) peptide = seq[start:end] # Optional: Verify the target residue matches to catch position errors if seq[target_pos - 1] != target_residue: print(f"Warning: Residue at position {target_pos} in {uniprot_id} is {seq[target_pos-1]}, not {target_residue}") return peptide # Replace this with your actual input data input_data = [("Q7TQ48_", "S", 442)] for item in input_data: uniprot_id, res, pos = item peptide = extract_13mer(uniprot_id, res, pos) print(f"{uniprot_id}: {peptide}")Pro tip: If you have a large list, add a small delay between requests to avoid hitting Uniprot’s rate limits.
For R users, the UniprotR package simplifies fetching sequences, and basic string manipulation will handle the peptide extraction:
Install the required package:
install.packages("UniprotR")Use this script:
library(UniprotR) extract_13mer <- function(uniprot_id, target_residue, target_pos) { # Clean isoform suffix if present clean_id <- strsplit(uniprot_id, "_")[[1]][1] # Fetch the protein sequence seq_df <- GetProteinSequence(clean_id, format = "fasta") seq <- as.character(seq_df$Sequence) # Calculate bounds (both Uniprot and R use 1-based indexing) start <- max(1, target_pos - 6) end <- min(nchar(seq), target_pos + 6) peptide <- substr(seq, start, end) # Verify residue match to catch errors if (substr(seq, target_pos, target_pos) != target_residue) { warning(paste0("Residue mismatch in ", uniprot_id, ": expected ", target_residue, ", got ", substr(seq, target_pos, target_pos))) } return(peptide) } # Replace with your actual input data input_data <- data.frame( UniprotID = c("Q7TQ48_"), TargetResidue = c("S"), TargetPos = c(442) ) # Process all entries input_data$Peptide <- mapply(extract_13mer, input_data$UniprotID, input_data$TargetResidue, input_data$TargetPos) print(input_data)
If you prefer not to code, here’s a workflow using Uniprot’s built-in tools plus basic spreadsheet functions:
Batch fetch sequences:
- Use Uniprot’s batch retrieval feature to upload all your Uniprot IDs (strip isoform suffixes if needed), then export the results as a CSV/Excel file that includes the full protein sequences.
Extract peptides with spreadsheet formulas:
- In your spreadsheet, add a new column for the peptide. Assuming your sequence is in column B and target position in column C, use this formula (works for Excel/Google Sheets):
This automatically handles edge cases where the target position is within the first 6 or last 6 residues of the sequence.=MID(B2, MAX(1, C2-6), MIN(13, C2+6 - MAX(1, C2-6) + 1))
- In your spreadsheet, add a new column for the peptide. Assuming your sequence is in column B and target position in column C, use this formula (works for Excel/Google Sheets):
Verify residues:
- Add a column to cross-check the target residue:
Compare this to your target residue column to catch any position mismatches.=MID(B2, C2, 1)
- Add a column to cross-check the target residue:
内容的提问来源于stack exchange,提问作者Emanuele Loro

