Python嵌套字典优化:实现词频按句子统计的更简洁方案
Great question! Let's walk through several cleaner approaches to achieve your goal—using the setdefault method and nested defaultdict (as you mentioned), plus an optimized approach with Counter that might save you some cycles.
1. 嵌套 defaultdict(最推荐)
The collections.defaultdict was made exactly for scenarios like this, where you want auto-initialized nested structures. Here's how to adapt it to your needs:
from collections import defaultdict # 外层字典:键是单词,值是另一个defaultdict(键是句子索引,值是次数) dfm = defaultdict(lambda: defaultdict(int)) for idx, sentence in enumerate(sentences): for word in sentence: # 直接累加,不存在的键会自动初始化为0 dfm[word][idx] += 1
为什么这管用?
- 外层
defaultdict(lambda: defaultdict(int))会自动为任何不存在的单词创建一个内层的defaultdict(int)。 - 内层
defaultdict(int)会自动把不存在的句子索引对应的初始值设为0,所以你只需要直接+=1就行,完全省去了手动判断键是否存在的代码。
如果之后需要把它转换成普通字典(比如要序列化或避免后续默认值行为),可以这样做:
# 转换成普通嵌套字典 regular_dfm = {word: dict(counts) for word, counts in dfm.items()}
2. 使用 setdefault(无需导入模块)
If you prefer not to use collections, Python's built-in dict.setdefault() method lets you achieve the same result with minimal extra code:
dfm = {} for idx, sentence in enumerate(sentences): for word in sentence: # 为当前单词创建空字典(如果不存在的话) dfm.setdefault(word, {}) # 为当前句子索引设置初始值0(如果不存在的话),然后累加 dfm[word][idx] = dfm[word].get(idx, 0) + 1
Or even more compact, chaining setdefault calls:
dfm = {} for idx, sentence in enumerate(sentences): for word in sentence: # 链式调用:先确保单词对应的字典存在,再确保句子索引的初始值为0 dfm.setdefault(word, {}).setdefault(idx, 0) dfm[word][idx] += 1
解释
setdefault(key, default)checks ifkeyexists in the dictionary: if yes, it returns the value; if no, it addskey: defaultto the dictionary and returnsdefault.- This eliminates the need for
if word not in dfm:andif idx not in dfm[word]:checks.
3. 优化版:用 Counter 统计单句词频
For sentences with many repeated words, using collections.Counter to count word frequencies per sentence first can be more efficient (since you avoid looping over duplicate words multiple times):
from collections import Counter, defaultdict dfm = defaultdict(dict) for idx, sentence in enumerate(sentences): # 先统计当前句子的所有词频 sentence_counts = Counter(sentence) # 把结果合并到总字典中 for word, count in sentence_counts.items(): dfm[word][idx] = count
This approach is cleaner when dealing with long sentences, as it reduces the number of iterations over individual words.
对比各方案
| 方案 | 优点 | 缺点 |
|---|---|---|
嵌套 defaultdict | 代码最简洁,可读性最高,编写最快 | 需要导入collections模块 |
setdefault | 无需导入模块,轻量灵活 | 代码稍长,嵌套调用略繁琐 |
Counter+defaultdict | 处理重复单词多的句子效率更高,逻辑清晰 | 多一步统计,适合特定场景 |
内容的提问来源于stack exchange,提问作者Pablo Ruiz Ruiz

