Python 3:如何生成去除子串单词的词频字典?
我来帮你搞定这个问题!你的需求其实分两步:先统计每个单词的出现频率,再剔除那些是其他单词子串的单词。咱们一步步来实现:
第一步:清理数据并统计频率
首先注意到你的输入列表里每个单词后面都带空格(比如'goon '),这会导致统计时把'goon '和'goon'当成不同单词,所以第一步必须先清理这些空白字符。用Python标准库的Counter来统计频率非常高效:
from collections import Counter # 原始输入列表 word_list = ['goon ', 'goonk ', 'goon ', 'goonj ', 'w ', 'wo ', 'wor ', 'world ', 'world '] # 清理每个单词的前后空白,再统计频率 cleaned_words = [word.strip() for word in word_list] frequency_counter = Counter(cleaned_words)
这一步后,frequency_counter会得到:Counter({'goon': 2, 'world': 2, 'goonk': 1, 'goonj': 1, 'w': 1, 'wo': 1, 'wor': 1})
第二步:筛选非子串单词
接下来要找出那些不是任何其他单词子串的单词。核心逻辑是:遍历每个唯一单词,检查是否存在另一个不同的单词包含它作为子串,若存在则剔除,否则保留。
# 获取所有唯一单词 unique_words = list(frequency_counter.keys()) # 筛选符合条件的单词 valid_words = [] for word in unique_words: # 检查是否存在其他单词(排除自身)包含当前单词 is_substring = any(other_word != word and word in other_word for other_word in unique_words) if not is_substring: valid_words.append(word) # 构建最终结果字典 result_dict = {word: frequency_counter[word] for word in valid_words} print(result_dict)
运行这段代码,输出正好是你想要的:{'goonj': 1, 'world': 2, 'goonk': 1}
可选优化:提升效率
如果你的单词列表规模很大,上面的O(n²)复杂度可能有点慢。可以先把单词按长度从长到短排序——长单词不可能是短单词的子串,所以只需要检查前面更长的单词是否包含当前单词,减少比较次数:
# 按单词长度降序排序 unique_words_sorted = sorted(unique_words, key=lambda x: -len(x)) valid_words = [] for i, word in enumerate(unique_words_sorted): # 只对比前面更长的单词 is_substring = any(word in other_word for other_word in unique_words_sorted[:i]) if not is_substring: valid_words.append(word) result_dict = {word: frequency_counter[word] for word in valid_words}
内容的提问来源于stack exchange,提问作者Lakshya Aggarwal
相关产品推荐
相关产品推荐

