如何优化基于字典统计文件单词出现次数的Python代码?
Hey there! Your current code gets the job done, but storing every word in that flat_list before counting is definitely inefficient—you’re wasting extra memory holding all those words, and then adding a second loop just to tally up counts. Let’s fix this by updating the dictionary as we read each word instead of using that intermediate list.
Optimized Code
def create_dictionary(filename): word_counts = {} with open(filename, "r") as fin: for line in fin: # Process each word immediately after splitting the line for word in line.split(): # Use dict.get() to safely increment counts without if/else checks word_counts[word] = word_counts.get(word, 0) + 1 return word_counts
What’s Improved & Why
- No extra memory for a list: We skip storing every word in
flat_listentirely—this is a huge win for large files, where that list could balloon to take up tons of unnecessary space. - Single-pass processing: We count words the moment we read them, eliminating the need for a second loop over the list. This cuts down on runtime by doing all the work in one go.
- Cleaner logic: Using
word_counts.get(word, 0) + 1removes the clunkyif/elsecheck—it’s a more Pythonic way to handle incrementing values for keys that might not exist yet.
Test It With Your Example
If your file has this content:
i go to the market to buy some things to eat and drink because i want to eat and drink
The optimized function will return exactly the expected result:
{'i': 2, 'go': 1, 'to': 4, 'the': 1, 'market': 1, 'buy': 1, 'some': 1, 'things': 1, 'eat': 2, 'and': 2, 'drink': 2, 'because': 1, 'want': 1}
Bonus: Even More Pythonic with collections.defaultdict
If you’re okay with using Python’s standard library, defaultdict can simplify the code even further by handling default values for new keys automatically:
from collections import defaultdict def create_dictionary(filename): word_counts = defaultdict(int) with open(filename, "r") as fin: for line in fin: for word in line.split(): word_counts[word] += 1 # Convert to a regular dict if you need the exact return type as your original code return dict(word_counts)
This works because defaultdict(int) sets any new key’s value to 0 automatically when you first access it, so you can just increment directly without using get().
内容的提问来源于stack exchange,提问作者gmarchale

