代码调试求助:复合关键词匹配函数未返回预期结果
我的代码哪里出错了?
代码目标
判断所有指定短语是否存在于输入字符串中。如果所有短语都能匹配(以distance作为容错范围)则返回True,否则返回False。示例:
- input = 'can i go to the bathroom in the morning'
- phrases = ['can go', 'bathroom morning']
- 当
distance为1时无法匹配,因为'bathroom'和'morning'之间有2个单词 - 当
distance为2时,'bathroom in the morning'被视为有效匹配短语
预期输出
input = "can i go to the bathroom in the morning" phrases = ['can go', 'bathroom morning'] distance = 2 print('Output',get_compound_keyword_match(input, phrases, distance))
输出:Output True
我的代码
def get_compound_keyword_match(input: str, phrases: list, distance: int) -> bool: if not distance: # We have no leeway for a match. if all(phrase in input for phrase in phrases): return True keywords = input.split() for phrase in phrases: phrase_matched = False ck_words = phrase.split() first_word_matches = [ i for i, x in enumerate(keywords) if x == ck_words[0] ] print('first word matches', first_word_matches) if not first_word_matches: return False for first_word_match in first_word_matches: old_match_index = first_word_match matched = False for i in range(0, len(ck_words)): try: match_index = keywords.index(ck_words[i]) if match_index - old_match_index > (distance + 1): matched = False old_match_index = match_index except ValueError: print('value error false') matched = False if matched: phrase_matched = True break if not phrase_matched: print('phrase_matched false') return False return True if __name__ == "__main__": input = "can i go to the bathroom in the morning" phrases = ['can go', 'bathroom morning'] distance = 2 print('Output',get_compound_keyword_match(input, phrases, distance))
请帮忙排查我的代码为何无法返回预期的True结果?
代码问题排查
你的代码存在三个关键问题,直接导致无法返回预期结果:
keywords.index()的使用逻辑错误index()方法只会返回第一个匹配项的索引,不会从当前old_match_index的位置往后查找。比如处理'bathroom morning'时,ck_words[1]是'morning',keywords.index('morning')会直接返回最后一个索引,而不是从bathroom的索引(5)之后查找,这会让距离判断完全失效。matched变量从未被设为True
你初始化matched = False后,只有在不满足距离条件或找不到单词时才设置matched = False,但没有任何逻辑会将matched改为True。哪怕所有单词都符合要求,matched也一直是False,导致phrase_matched无法被激活,最终返回False。循环遍历逻辑冗余且错误
处理短语中的单词时,你从索引0开始循环,但第一个单词的位置已经通过first_word_match确定,不需要再重新查找。这种重复查找不仅无意义,还会打乱后续的距离判断逻辑。
修复后的代码示例
def get_compound_keyword_match(input: str, phrases: list, distance: int) -> bool: if not distance: return all(phrase in input for phrase in phrases) keywords = input.split() for phrase in phrases: phrase_matched = False ck_words = phrase.split() # 获取第一个单词的所有出现位置 first_word_indices = [i for i, word in enumerate(keywords) if word == ck_words[0]] if not first_word_indices: return False for start_idx in first_word_indices: current_idx = start_idx matched_all = True # 从第二个单词开始检查,基于前一个单词的位置限定查找范围 for word in ck_words[1:]: found = False # 最多往后查找distance+1个位置(包含容错的单词数) for i in range(current_idx + 1, min(current_idx + distance + 2, len(keywords))): if keywords[i] == word: current_idx = i found = True break if not found: matched_all = False break if matched_all: phrase_matched = True break if not phrase_matched: return False return True if __name__ == "__main__": input_str = "can i go to the bathroom in the morning" phrases = ['can go', 'bathroom morning'] distance = 2 print('Output', get_compound_keyword_match(input_str, phrases, distance))
修复说明
- 替换
index()为从当前位置向后遍历查找,确保找到的是当前单词之后的匹配项 - 新增
matched_all变量,只有当所有单词都在允许的距离内找到时才标记为True - 调整循环逻辑,从短语的第二个单词开始检查,基于前一个单词的位置限定查找范围
- 优化变量名(如
input改为input_str,避免和内置函数冲突)
内容的提问来源于stack exchange,提问作者user7104332
相关产品推荐
相关产品推荐

