Python无模块导入:统计文本文件各长度单词的出现次数
Solution to Count Word Lengths into Ordered Dictionary
Got it, let's fix this step by step. You're on the right track opening the file—here's how to build out that dictify function to count word lengths and get the ordered dictionary you want, no modules needed:
Complete Function Code
def dictify(book_name): word_length_counts = {} # Use `with` statement to safely handle file I/O (auto-closes file) with open(book_name, 'r') as infile: # Process file line by line to save memory for large texts for line in infile: # Split line into individual words (splits on all whitespace) words = line.split() for word in words: # Optional: Clean punctuation to count only alphabetic characters # Uncomment below if you need pure word length (no punctuation) # cleaned_word = ''.join(char for char in word if char.isalpha()) # if not cleaned_word: # continue # Skip empty strings from all-punctuation "words" # word_length = len(cleaned_word) # Use raw word length if you don't need to remove punctuation word_length = len(word) # Update the count dictionary if word_length in word_length_counts: word_length_counts[word_length] += 1 else: word_length_counts[word_length] = 1 # Sort the dictionary by word length (key) to get ordered output sorted_counts = dict(sorted(word_length_counts.items())) return sorted_counts
Key Explanations
- Safe File Handling: The
withstatement ensures the file is automatically closed after processing, preventing resource leaks that can happen with manualclose()calls. - Memory Efficiency: Processing the file line by line avoids loading the entire book into memory, which is crucial for large text files.
- Word Processing:
line.split()breaks each line into a list of words using whitespace as the delimiter.- The optional punctuation cleaning step filters out non-alphabetic characters (like commas, periods, or exclamation marks) if you want to count the length of the actual word, not the word plus attached punctuation.
- Count Logic: For each word length, we check if it's already a key in our dictionary. If it is, we increment the count; if not, we add it to the dictionary with an initial count of 1.
- Ordered Output: Using
sorted(word_length_counts.items())sorts the dictionary entries by word length (the key) in ascending order, then converts the sorted list back to a dictionary to get the{1: 1482, 2: 2042...}format you need.
This function will return exactly the structured, ordered count dictionary you're aiming for. Feel free to tweak the punctuation handling if your text has specific edge cases!
内容的提问来源于stack exchange,提问作者anonymous
相关产品推荐
相关产品推荐

