Python中使用for循环向列表添加元素出现大量重复的问题排查
筛选包含指定字符的单词并避免重复
问题描述
我是一名Python初学者,想要生成一个新列表discards,仅包含wordssplit中含有alphabet列表内特定字符的单词。尝试用for循环实现时,出现了大量重复单词的问题。
原代码:
lowercase = text.lower() wordssplit = lowercase.split() alphabet = ["a", "b", "z", "h"] discards = [] for x in wordssplit: if x not in discards: for p in alphabet: if p in x: discards.append(x) print(discards, end='') txt("could you please show me the code")
输出:
[] [] ['please'] ['please', 'show'] ['please', 'show'] ['please', 'show', 'the'] ['please', 'show', 'the']
原因分析
- 重复添加的核心原因:内层循环遍历
alphabet的每个字符时,只要字符存在于单词中就执行一次append(x)。如果一个单词包含多个alphabet中的字符,就会被多次添加到discards里。 - 判断时机错误:
if x not in discards的判断只在外层循环开始时执行一次,内层循环中即使已经把x添加到discards,后续匹配到其他字符时仍会重复添加。
解决方法
方法一:用any()简化判断,避免重复添加
通过any()函数一次性判断单词是否包含alphabet中的任意字符,再检查是否已在列表中,符合条件才添加:
text = "could you please show me the code" lowercase = text.lower() wordssplit = lowercase.split() alphabet = ["a", "b", "z", "h"] discards = [] for x in wordssplit: # 先判断单词是否符合字符条件,再判断是否已存在 if any(p in x for p in alphabet) and x not in discards: discards.append(x) print(discards)
输出:
[] [] ['please'] ['please', 'show'] ['please', 'show'] ['please', 'show', 'the'] ['please', 'show', 'the']
注:输出中的重复打印是因为每个单词循环都会执行一次print,如果不需要逐次打印,将print(discards)移到循环外即可。
方法二:用集合优化去重效率(推荐)
列表的x in discards查找效率较低(O(n)),用集合记录已添加的单词,既能避免重复,又能提升查找速度:
text = "could you please show me the code" lowercase = text.lower() wordssplit = lowercase.split() alphabet = ["a", "b", "z", "h"] discards = [] seen = set() for x in wordssplit: if any(p in x for p in alphabet) and x not in seen: discards.append(x) seen.add(x) print(discards) # 输出:['please', 'show', 'the']
这种方法既能保持单词在原列表中的顺序,又能高效去重,适合处理大规模数据。
内容的提问来源于stack exchange,提问作者Aresst
相关产品推荐
相关产品推荐

