Python实现句子变位词检测并删除非重复变位词对中一个的方法
问题原因分析
- 原代码仅对按长度排序后的相邻单词做变位词校验,非相邻的变位词不会被匹配到
- 原代码遗漏了相同重复单词需全部保留的判断逻辑
- 原代码存在冗余调用
verify_anagram的问题,第一次调用的返回值未被使用,额外消耗性能
修复方案
核心逻辑是提前记录已经保留的单词,遍历所有单词逐个判断,不需要提前按长度排序,优先保留先出现的变位词,完全复用你已实现的verify_anagram函数即可:
- 相同单词直接保留,无论是否重复
- 新单词和所有已保留的同长度单词做变位词校验,只要存在互为变位词的已保留单词就删除当前词
- 新的变位词直接保留,加入已保留列表中
修改后代码
def correct_doc(doc): # 校验输入合法性 if not isinstance(doc, str): raise ValueError() if not (all(c.isalpha() or c.isspace() for c in doc) and doc.count(' ') == 0): raise ValueError() words = doc.split() kept_words = [] seen_exact_words = set() for word in words: # 相同单词直接保留,符合重复相同单词全留的要求 if word in seen_exact_words: kept_words.append(word) continue # 校验当前词是否和已保留的单词互为变位词 has_anagram_exist = False for kept_word in kept_words: # 长度不同不可能是变位词,提前跳过减少校验次数 if len(kept_word) != len(word): continue if verify_anagram(kept_word, word): has_anagram_exist = True break if has_anagram_exist: continue # 新的变位词加入保留列表 kept_words.append(word) seen_exact_words.add(word) return ' '.join(kept_words)
效果验证
- 输入示例句子:
John was going for a race but he has to take care of his son,输出为John was going for a race but he has to take of his son,后出现的变位词care被删除,符合预期。 - 输入重复相同单词的句子:
data analysis needs data support,输出和输入完全一致,两个data都被保留,符合预期。
内容的提问来源于stack exchange,提问作者user17294101
相关产品推荐
相关产品推荐

