You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CS50 DNA序列匹配程序始终返回‘无匹配’的逻辑错误排查

CS50 DNA匹配程序无匹配问题修复

我正在完成CS50的DNA匹配作业,程序需要根据包含姓名和STR重复次数的CSV数据库,与给定DNA序列匹配。已知longest_match函数能正确返回STR最长连续重复次数,且调试确认变量值符合预期,但程序始终返回“无匹配”,本应匹配到数据库中的Bob。


代码

import csv
import sys

def main():

    # Check for command-line usage
    if len(sys.argv) != 3:
        print("Usage: python python.py ____.csv ___.csv")
        sys.exit(1)
    # Read database file into a variable
    databaseName = sys.argv[1]
    DNASequence = sys.argv[2]

    with open(databaseName, 'r') as file:
        reader = csv.DictReader(file)
        people = []
        subSeq = reader.fieldnames


        for row in reader:
            people.append(row)

    # Read DNA sequence file into a variable
    with open(DNASequence, 'r') as file:
        sequence = file.read()
    # Find longest match of each STR in DNA sequence
    runs = []

    for i in range(1, len(subSeq)):
        sub = {subSeq[i]: longest_match(sequence, subSeq[i])}
        runs.append(sub)

    # Really proud of this one ^^^
    # Check database for matching profiles
    for dict1 in people:
        match = True

        for i in range(1, len(subSeq)):
            current_key = subSeq[i]

            if dict1[current_key] != runs[i - 1][current_key]:
                match = False

        if match == True:
            print(f"Match found: {dict1[subSeq[0]]}")

    if match == False:
        print("No match found.")

    return

def longest_match(sequence, subsequence):
    """Returns length of longest run of subsequence in sequence."""

    # Initialize variables
    longest_run = 0
    subsequence_length = len(subsequence)
    sequence_length = len(sequence)

    # Check each character in sequence for most consecutive runs of subsequence
    for i in range(sequence_length):

        # Initialize count of consecutive runs
        count = 0

        # Check for a subsequence match in a "substring" (a subset of characters) within sequence
        # If a match, move substring to next potential match in sequence
        # Continue moving substring and checking for matches until out of consecutive matches
        while True:

            # Adjust substring start and end
            start = i + count * subsequence_length
            end = start + subsequence_length

            # If there is a match in the substring
            if sequence[start:end] == subsequence:
                count += 1

            # If there is no match in the substring
            else:
                break

        # Update most consecutive matches found
        longest_run = max(longest_run, count)

    # After checking for runs at each character in seqeuence, return longest run found
    return longest_run


main()

CSV样例

name,AGATC,AATG,TATC
Alice,2,8,3
Bob,4,1,5
Charlie,3,2,5

测试DNA序列

AAGGTAAGTTTAGAATATAAAAGGTGAGTTAAATAGAATAGGTTAAAATTAAAGGAGATCAGATCAGATCAGATCTATCTATCTATCTATCTATCAGAAAAGAGTAAATAGTTAAAGAGTAAGATATTGAATTAATGGAAAATATTGTTGGGGAAAGGAGGGATAGAAGG

问题原因

  1. 类型不匹配:从CSV读取的STR重复次数是字符串类型(比如Bob的AGATC值是"4"),而longest_match返回的是整数类型(比如4),直接比较字符串和整数会导致判断为不相等,触发match = False。
  2. 最终输出逻辑错误:match变量在循环每个人员时都会被重置,若最后一个人员不匹配,即使前面已经找到匹配(比如Bob),程序仍会输出“无匹配”。

修复方案

1. 修正类型比较

将CSV中读取的字符串类型转换为整数,再与runs中的整数比较:

if int(dict1[current_key]) != runs[i - 1][current_key]:
    match = False

2. 修正最终输出逻辑

新增一个found变量记录是否找到匹配,避免循环覆盖导致的错误,同时可以在找到匹配后提前终止循环优化效率:

# 初始化匹配标记
found = False

for dict1 in people:
    match = True

    for i in range(1, len(subSeq)):
        current_key = subSeq[i]

        if int(dict1[current_key]) != runs[i - 1][current_key]:
            match = False
            break  # 一旦不匹配,无需继续比较当前人员的其他STR

    if match:
        print(f"Match found: {dict1[subSeq[0]]}")
        found = True
        # 可选:找到第一个匹配后直接退出程序
        # sys.exit(0)

if not found:
    print("No match found.")

优化后的完整main函数

def main():

    # Check for command-line usage
    if len(sys.argv) != 3:
        print("Usage: python python.py database.csv sequence.txt")
        sys.exit(1)
    # Read database file into a variable
    databaseName = sys.argv[1]
    DNASequence = sys.argv[2]

    with open(databaseName, 'r') as file:
        reader = csv.DictReader(file)
        people = []
        subSeq = reader.fieldnames

        for row in reader:
            people.append(row)

    # Read DNA sequence file into a variable
    with open(DNASequence, 'r') as file:
        sequence = file.read()
    # Find longest match of each STR in DNA sequence
    runs = []

    for i in range(1, len(subSeq)):
        sub = {subSeq[i]: longest_match(sequence, subSeq[i])}
        runs.append(sub)

    # Check database for matching profiles
    found = False
    for dict1 in people:
        match = True

        for i in range(1, len(subSeq)):
            current_key = subSeq[i]

            if int(dict1[current_key]) != runs[i - 1][current_key]:
                match = False
                break

        if match:
            print(f"Match found: {dict1[subSeq[0]]}")
            found = True
            # 可选:找到第一个匹配后直接退出程序
            # sys.exit(0)

    if not found:
        print("No match found.")

    return

内容的提问来源于stack exchange,提问作者Charlie Webster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 15:51:02