不使用内置函数或导入库的Python文本文件单词计数作业难题
Hey there! Let's break down what's going wrong with your code and fix it step by step.
What's Causing the Issue?
Your current code has a critical problem in the else block: when you encounter a word that's already in compare_list, you're looping through the entire master_list every single time to increment the count. This means if a word appears 5 times, you'll end up adding 5 to its count multiple times (once for each occurrence), leading to wildly inflated and incorrect numbers.
The Fix: Targeted Count Increment
Instead of re-scanning the entire list every time you find a duplicate word, just find the index of the existing word in compare_list and increment the corresponding position in count_list directly. Here's the revised code:
# Most common word counter fh = open("romeo.txt") master_list = fh.read().split() fh.close() # Don't forget to close the file! compare_list = [] count_list = [] for word in master_list: if word not in compare_list: compare_list.append(word) count_list.append(1) else: # Find the index of the existing word and increment its count word_index = compare_list.index(word) count_list[word_index] += 1 # Now find the most common word max_count = max(count_list) most_common_index = count_list.index(max_count) most_common_word = compare_list[most_common_index] print(f"Total words: {len(master_list)}") print(f"Unique words: {len(compare_list)}") print(f"Counts per word: {list(zip(compare_list, count_list))}") print(f"The most common word is '{most_common_word}' with {max_count} occurrences.")
Key Improvements:
- Fixed Counting Logic: We use
compare_list.index(word)to directly find where the duplicate word lives, then increment its count by 1 (no more re-counting the entire list). - Proper File Handling: Added
fh.close()to properly close the text file after reading it—this is good practice to avoid wasting system resources. - Final Result Calculation: Added code to actually find and print the most common word, which was missing from your original code.
- Readable Output: Used
zip()to pair words with their counts for clearer, easier-to-interpret printing.
Optional: Case Insensitivity
If your assignment considers "Romeo" and "romeo" as the same word, you can normalize all words to lowercase (or uppercase) before processing to avoid skewed results:
# Modify the loop to use lowercase words for word in master_list: normalized_word = word.lower() if normalized_word not in compare_list: compare_list.append(normalized_word) count_list.append(1) else: word_index = compare_list.index(normalized_word) count_list[word_index] += 1
内容的提问来源于stack exchange,提问作者timothy johnsonII

