You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3:如何生成去除子串单词的词频字典?

我来帮你搞定这个问题!你的需求其实分两步:先统计每个单词的出现频率,再剔除那些是其他单词子串的单词。咱们一步步来实现:

第一步:清理数据并统计频率

首先注意到你的输入列表里每个单词后面都带空格(比如'goon '),这会导致统计时把'goon '和'goon'当成不同单词,所以第一步必须先清理这些空白字符。用Python标准库的Counter来统计频率非常高效:

from collections import Counter

# 原始输入列表
word_list = ['goon ', 'goonk ', 'goon ', 'goonj ', 'w ', 'wo ', 'wor ', 'world ', 'world ']

# 清理每个单词的前后空白,再统计频率
cleaned_words = [word.strip() for word in word_list]
frequency_counter = Counter(cleaned_words)

这一步后,frequency_counter会得到:
Counter({'goon': 2, 'world': 2, 'goonk': 1, 'goonj': 1, 'w': 1, 'wo': 1, 'wor': 1})

第二步:筛选非子串单词

接下来要找出那些不是任何其他单词子串的单词。核心逻辑是:遍历每个唯一单词,检查是否存在另一个不同的单词包含它作为子串,若存在则剔除,否则保留。

# 获取所有唯一单词
unique_words = list(frequency_counter.keys())

# 筛选符合条件的单词
valid_words = []
for word in unique_words:
    # 检查是否存在其他单词(排除自身)包含当前单词
    is_substring = any(other_word != word and word in other_word for other_word in unique_words)
    if not is_substring:
        valid_words.append(word)

# 构建最终结果字典
result_dict = {word: frequency_counter[word] for word in valid_words}
print(result_dict)

运行这段代码,输出正好是你想要的:
{'goonj': 1, 'world': 2, 'goonk': 1}

可选优化:提升效率

如果你的单词列表规模很大,上面的O(n²)复杂度可能有点慢。可以先把单词按长度从长到短排序——长单词不可能是短单词的子串,所以只需要检查前面更长的单词是否包含当前单词,减少比较次数:

# 按单词长度降序排序
unique_words_sorted = sorted(unique_words, key=lambda x: -len(x))
valid_words = []
for i, word in enumerate(unique_words_sorted):
    # 只对比前面更长的单词
    is_substring = any(word in other_word for other_word in unique_words_sorted[:i])
    if not is_substring:
        valid_words.append(word)

result_dict = {word: frequency_counter[word] for word in valid_words}

内容的提问来源于stack exchange,提问作者Lakshya Aggarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:13:15