如何为胰蛋白酶酶解肽段列表设置连续索引以定位指定氨基酸
Got it, let's work through this problem to automate the peptide position tracking and amino acid lookup—no more manual counting, even for thousand-residue sequences!
Step 1: Map Each Peptide to Its Start/End Indices in the Original Sequence
First, we'll iterate through your peptide list and calculate the 0-based start and end indices for each segment, using the previous peptide's end index as the next one's start. This builds a full position map tied to the original sequence.
# Your original sequence and tryptic peptides string_sequence = 'MLSPDLPDSAWNTRLLCRVMLCLLGAGSVAAGVIQSPRHLIKEKRETATLKCYPIPRHDTVYWYQQGPGQDPQFLISFYEKMQSDKGSIPDRFSAQQFSDYHSELNMSSLELGDSALYFCASSL' list_sequence = ['MLSPDLPDSAWNTR', 'LLCR', 'VMLCLLGAGSVAAGVIQSPR', 'HLIK', 'EK', 'R', 'ETATLK', 'CYPIPR', 'HDTVYWYQQGPGQDPQFLISFYEK', 'MQSDK', 'GSIPDR', 'FSAQQFSDYHSELNMSSLELGDSALYFCASSL'] # Store each peptide with its start/end indices (0-based in original sequence) peptide_position_map = [] current_start = 0 for peptide in list_sequence: peptide_len = len(peptide) current_end = current_start + peptide_len # Append tuple: (peptide string, start index, end index) peptide_position_map.append( (peptide, current_start, current_end) ) # Update start for the next peptide current_start = current_end
Step 2: Locate the Target Amino Acid from the Original Sequence
Next, we'll write a helper function to take an original sequence position (either 1-based or 0-based) and find exactly which peptide it's in, plus its position within that peptide.
For your example, the original sequence's 18th amino acid (1-based) translates to index 17 in 0-based notation (since programming languages typically use 0-based indexing).
def find_amino_acid_position(target_original_idx, position_map): # target_original_idx: 0-based index from the original sequence for peptide_index, (peptide, start, end) in enumerate(position_map): # Check if the target index falls within this peptide's range if start <= target_original_idx < end: # Calculate position inside the peptide (both 0-based and 1-based) internal_0based = target_original_idx - start internal_1based = internal_0based + 1 return { 'peptide_list_index': peptide_index, 'peptide_sequence': peptide, 'original_start': start, 'original_end': end, 'position_in_peptide_0based': internal_0based, 'position_in_peptide_1based': internal_1based } # Return None if the index is out of bounds return None # Example: Look up the 18th amino acid (1-based) in the original sequence target_1based = 18 target_0based = target_1based - 1 result = find_amino_acid_position(target_0based, peptide_position_map) if result: print(f"Original position {target_1based} (0-based {target_0based}) found:") print(f"- Peptide index in list: {result['peptide_list_index']}") print(f"- Peptide sequence: {result['peptide_sequence']}") print(f"- Position within peptide (1-based): {result['position_in_peptide_1based']}") else: print("Target index is outside the original sequence's length.")
When you run this, you'll get output confirming that the 18th amino acid (V) is the first residue (1-based) in the third peptide (index 2 in the list)—which matches manual verification.
Step 3: Verify Your Peptide List is Correct
To make sure your tryptic digest list doesn't have missing or extra residues (critical for large sequences), add this quick check:
total_peptide_length = sum(len(p) for p in list_sequence) original_length = len(string_sequence) print(f"Total peptide length: {total_peptide_length}") print(f"Original sequence length: {original_length}") print(f"Peptide list matches original sequence: {total_peptide_length == original_length}")
This should print True if your peptide list is a complete, non-overlapping split of the original sequence.
内容的提问来源于stack exchange,提问作者menbar

