基因组序列模式匹配代码故障排查:仅返回首次匹配位置而非全部匹配位置
问题分析与修复方案
问题根源
你代码里的核心bug出在匹配位置的获取逻辑上:每次找到匹配时,你调用了Text.index(Pattern)来添加位置,但这个方法的特性是仅返回模式在文本中第一次出现的索引,不管循环遍历到哪个位置,它都会重复返回同一个起始值,这就导致所有匹配位置都被替换成了首次匹配的索引。
修复后的代码
# open the file with the original sequence myfile = open('Vibrio_cholerae.txt') # set the file to the variable Text to read and scan Text = myfile.read() # insert the pattern Pattern = "TAATGGCT" PatternLocations = [] def PatternCount(Text, Pattern): count = 0 pattern_len = len(Pattern) text_len = len(Text) for i in range(text_len - pattern_len + 1): if Text[i:i+pattern_len] == Pattern: count += 1 # 直接使用当前循环的i作为匹配起始位置,而不是index() PatternLocations.append(i) return count # print the result of calling PatternCount on Text and Pattern. print(f"Number of times the Pattern is repeated: {PatternCount(Text, Pattern)} time(s).") print(f"List of Pattern locations: {PatternLocations}")
关键修正说明
- 把
PatternLocations.append(Text.index(Pattern))替换成PatternLocations.append(i):循环变量i就是当前匹配片段的起始索引,这正是我们需要记录的每个匹配的真实位置。 - 额外优化:提前计算
pattern_len和text_len,避免在循环中重复计算长度,提升一点代码效率。
内容的提问来源于stack exchange,提问作者Nadeen NBY
相关产品推荐
相关产品推荐

