如何基于参考列表实现字符串精准匹配,避开正则及内容重叠问题
解决方案
你遇到的输出错误根源是短参考选项(比如示例中的No)是长参考选项的子串,短选项提前匹配到了长选项里的对应内容导致误判。我们可以通过长选项优先匹配+匹配后移除对应文本的逻辑解决,全程无需正则,也不需要处理分隔符。
实现代码
reference = ['I am not sure', "I don't know who to talk to", 'No', 'Not listed (specify)', 'I do not feel comfortable answering this question','Not really describing my situation'] answer = "I don't know who to talk to, Not listed (specify)" # 1. 参考列表按文本长度倒序排序,保留原始1开始的索引 sorted_reference = sorted(enumerate(reference, 1), key=lambda x: -len(x[1])) matched_indexes = [] current_answer = answer for origin_idx, choice_text in sorted_reference: if choice_text in current_answer: matched_indexes.append(origin_idx) # 匹配到后移除对应文本,避免后续短选项误匹配 current_answer = current_answer.replace(choice_text, '', 1) # 2. 按原始索引从小到大排序后拼接 result = ' '.join(map(str, sorted(matched_indexes))) print(result)
输出结果
2 4
方案说明
- 优先匹配长度更长的参考选项,从根源避免短选项匹配长选项子串的问题
- 匹配后直接从待处理字符串中删除已匹配内容,不会出现重复匹配、误匹配的情况
- 全程仅使用基础字符串操作,不依赖正则、不需要处理复杂分隔符,兼容存在特殊标点的场景
内容的提问来源于stack exchange,提问作者RubberDuckProducer
相关产品推荐
相关产品推荐

