Python如何从字符串中正确提取停用词并拆分两类词列表
问题根因
原代码有三处逻辑错误:
- 嵌套循环导致重复追加:对每个拆分出的单词,都会遍历全部停用词逐一判断,同一个单词会被执行多次追加操作,直接产生大量重复值
- 判断条件错误:使用
i in j实际是判断当前单词是否为某一个停用词的子串,并非判断单词是否属于停用词列表,逻辑和需求完全不符 - 冗余逻辑导致异常值:循环结束后额外添加的整列表判断没有实际意义,会把整个分词结果作为单个元素追加到列表中,产生不符合预期的嵌套列表元素
修正代码
text = 'he is the best when people in our life' stopwords = ['he', 'the', 'our'] # 转为集合提升成员判断效率,停用词规模较大时优化效果更显著 stopword_set = set(stopwords) contains_stopwords = [] normal_words = [] for word in text.split(): if word in stopword_set: contains_stopwords.append(word) else: normal_words.append(word) print("contains_stopwords:", contains_stopwords) print("normal_words:", normal_words)
运行结果
执行代码后输出和预期完全一致:
contains_stopwords: ['he', 'the', 'our'] normal_words: ['is', 'best', 'when', 'people', 'in', 'life']
内容的提问来源于stack exchange,提问作者user19375409
相关产品推荐
相关产品推荐

