CS50 DNA序列匹配程序始终返回‘无匹配’的逻辑错误排查
CS50 DNA匹配程序无匹配问题修复
我正在完成CS50的DNA匹配作业,程序需要根据包含姓名和STR重复次数的CSV数据库,与给定DNA序列匹配。已知longest_match函数能正确返回STR最长连续重复次数,且调试确认变量值符合预期,但程序始终返回“无匹配”,本应匹配到数据库中的Bob。
代码
import csv import sys def main(): # Check for command-line usage if len(sys.argv) != 3: print("Usage: python python.py ____.csv ___.csv") sys.exit(1) # Read database file into a variable databaseName = sys.argv[1] DNASequence = sys.argv[2] with open(databaseName, 'r') as file: reader = csv.DictReader(file) people = [] subSeq = reader.fieldnames for row in reader: people.append(row) # Read DNA sequence file into a variable with open(DNASequence, 'r') as file: sequence = file.read() # Find longest match of each STR in DNA sequence runs = [] for i in range(1, len(subSeq)): sub = {subSeq[i]: longest_match(sequence, subSeq[i])} runs.append(sub) # Really proud of this one ^^^ # Check database for matching profiles for dict1 in people: match = True for i in range(1, len(subSeq)): current_key = subSeq[i] if dict1[current_key] != runs[i - 1][current_key]: match = False if match == True: print(f"Match found: {dict1[subSeq[0]]}") if match == False: print("No match found.") return def longest_match(sequence, subsequence): """Returns length of longest run of subsequence in sequence.""" # Initialize variables longest_run = 0 subsequence_length = len(subsequence) sequence_length = len(sequence) # Check each character in sequence for most consecutive runs of subsequence for i in range(sequence_length): # Initialize count of consecutive runs count = 0 # Check for a subsequence match in a "substring" (a subset of characters) within sequence # If a match, move substring to next potential match in sequence # Continue moving substring and checking for matches until out of consecutive matches while True: # Adjust substring start and end start = i + count * subsequence_length end = start + subsequence_length # If there is a match in the substring if sequence[start:end] == subsequence: count += 1 # If there is no match in the substring else: break # Update most consecutive matches found longest_run = max(longest_run, count) # After checking for runs at each character in seqeuence, return longest run found return longest_run main()
CSV样例
name,AGATC,AATG,TATC Alice,2,8,3 Bob,4,1,5 Charlie,3,2,5
测试DNA序列
AAGGTAAGTTTAGAATATAAAAGGTGAGTTAAATAGAATAGGTTAAAATTAAAGGAGATCAGATCAGATCAGATCTATCTATCTATCTATCTATCAGAAAAGAGTAAATAGTTAAAGAGTAAGATATTGAATTAATGGAAAATATTGTTGGGGAAAGGAGGGATAGAAGG
问题原因
- 类型不匹配:从CSV读取的STR重复次数是字符串类型(比如Bob的
AGATC值是"4"),而longest_match返回的是整数类型(比如4),直接比较字符串和整数会导致判断为不相等,触发match = False。 - 最终输出逻辑错误:
match变量在循环每个人员时都会被重置,若最后一个人员不匹配,即使前面已经找到匹配(比如Bob),程序仍会输出“无匹配”。
修复方案
1. 修正类型比较
将CSV中读取的字符串类型转换为整数,再与runs中的整数比较:
if int(dict1[current_key]) != runs[i - 1][current_key]: match = False
2. 修正最终输出逻辑
新增一个found变量记录是否找到匹配,避免循环覆盖导致的错误,同时可以在找到匹配后提前终止循环优化效率:
# 初始化匹配标记 found = False for dict1 in people: match = True for i in range(1, len(subSeq)): current_key = subSeq[i] if int(dict1[current_key]) != runs[i - 1][current_key]: match = False break # 一旦不匹配,无需继续比较当前人员的其他STR if match: print(f"Match found: {dict1[subSeq[0]]}") found = True # 可选:找到第一个匹配后直接退出程序 # sys.exit(0) if not found: print("No match found.")
优化后的完整main函数
def main(): # Check for command-line usage if len(sys.argv) != 3: print("Usage: python python.py database.csv sequence.txt") sys.exit(1) # Read database file into a variable databaseName = sys.argv[1] DNASequence = sys.argv[2] with open(databaseName, 'r') as file: reader = csv.DictReader(file) people = [] subSeq = reader.fieldnames for row in reader: people.append(row) # Read DNA sequence file into a variable with open(DNASequence, 'r') as file: sequence = file.read() # Find longest match of each STR in DNA sequence runs = [] for i in range(1, len(subSeq)): sub = {subSeq[i]: longest_match(sequence, subSeq[i])} runs.append(sub) # Check database for matching profiles found = False for dict1 in people: match = True for i in range(1, len(subSeq)): current_key = subSeq[i] if int(dict1[current_key]) != runs[i - 1][current_key]: match = False break if match: print(f"Match found: {dict1[subSeq[0]]}") found = True # 可选:找到第一个匹配后直接退出程序 # sys.exit(0) if not found: print("No match found.") return
内容的提问来源于stack exchange,提问作者Charlie Webster
相关产品推荐
相关产品推荐

