求助:Python中sum结合sigma与pair的代码逻辑解析
Hey there! Let's break down that line of code step by step so it makes total sense. I know list comprehensions can look like a jumble at first, but once we unpack each part, it'll all click.
First, let's restate the full line for reference:
score = sum([-sigma if '-' in pair else scoring_matrix[pair] for pair in zip(v_aligned, w_aligned)])
This line calculates the total alignment score between two aligned sequences (v_aligned and w_aligned). Let's split it into digestible pieces:
1. The zip(v_aligned, w_aligned) part
The zip() function takes your two aligned strings and pairs up characters at the exact same position. For your example where:
v_aligned = LEASNT-LYw_aligned = MEA-N-LY-
zip() will generate a sequence of tuples like this:
('L','M'), ('E','E'), ('A','A'), ('S','-'), ('N','N'), ('T','-'), ('-','L'), ('L','Y'), ('Y','-')
Each tuple represents one position in the aligned sequences.
2. The List Comprehension (the part inside [])
This is the core logic that calculates a score for each position pair:
[-sigma if '-' in pair else scoring_matrix[pair] for pair in zip(...)]
It's a shorthand way to build a list of scores, one for each paired position:
- If the pair contains a
-(a gap in one of the sequences): We use-sigma(here, sigma=5, so this becomes-5). This is a gap penalty—we subtract points every time there's an unaligned position, since gaps indicate a less ideal match. - If the pair has no gaps: We look up the score for that amino acid pair in the
scoring_matrix(BLOSUM62 in your case). For example, the pair('E','E')would get a score of 4 (since BLOSUM62 rewards identical amino acids with a positive score), while a mismatched pair like('L','M')would get a lower (or negative) score based on how biochemically similar the amino acids are.
For your example, the generated list might look something like (using approximate BLOSUM62 values):
[-2, 4, 4, -5, 6, -5, -5, 2, -5]
3. The sum() Function
Finally, sum() takes all the individual position scores in that list and adds them together to get the total alignment score. This single number tells you how well the two sequences match up—higher scores mean a more biologically meaningful alignment.
To put it all in plain terms: We're going through every position in the aligned sequences, penalizing gaps, using the BLOSUM62 matrix for matches/mismatches, then adding up all those values to get the final measure of alignment quality.
内容的提问来源于stack exchange,提问作者Duha

