Python单词去重问题:不使用内置函数实现重复单词过滤
如何不使用内置函数实现Python重复单词过滤?
原代码如下:
input = 'hello world and practice makes perfect and hello world again' output=[] filter='' for word in input.split(): if word not in filter: filter = filter + word print(filter) else: continue output.append(filter) print(' '.join(output))
你期望的输出是:
hello world and practice makes perfect again
但实际得到的输出是:
hello helloworld helloworldand helloworldandpractice helloworldandpracticemakes helloworldandpracticemakesperfect helloworldandpracticemakesperfectagain
问题分析
你代码的核心错误在于用**字符串filter**来记录已出现的单词:
- 判断
word not in filter时,是检查单词是否是该字符串的子串,而非是否作为独立单词存在(比如如果有单词low,会被误判为在hello里); - 直接拼接字符串会把所有单词连在一起,无法区分独立单词,最终输出的是拼接后的长字符串而非空格分隔的单词列表。
修改后的代码
用列表存储已出现的单词,确保每个单词是独立的判断单元,就能解决问题:
input_str = 'hello world and practice makes perfect and hello world again' output = [] seen_words = [] # 用列表记录已经出现过的单词 for word in input_str.split(): if word not in seen_words: seen_words.append(word) output.append(word) # 输出最终结果 print(' '.join(output))
代码说明
seen_words列表专门用来记录已经处理过的单词,word not in seen_words能准确判断当前单词是否重复;- 每遇到不重复的单词,就同时添加到
seen_words(用于后续重复判断)和output(用于最终输出); - 最后用
' '.join(output)把列表转换成空格分隔的字符串,得到你期望的输出结果。
内容的提问来源于stack exchange,提问作者Hari
相关产品推荐
相关产品推荐

