技术问询:如何实现向莎士比亚单词列表添加单词并打印,以及从指定文本文件提取排序唯一单词?附错误代码修正需求
Fixing Your Shakespeare Unique Word List Code
Hey there! Let's work through fixing your code so it correctly extracts, deduplicates, sorts, and prints the unique words from romeo.txt. First, let's recap what we need this script to do:
- Read the file line by line
- Collect every unique word across all lines
- Sort those words alphabetically
- Print the final sorted list
What's Breaking the Original Code?
Your current code has a few critical logic missteps:
- Wrong list for tracking uniqueness: You're checking if a word is in
words(the list of words from the current line) and modifying that same list. This means you're only tracking unique words per line, not the entire file—plus, modifying a list while iterating over it can cause weird loops or duplicates. - Unnecessary file mode: Using
r+(read-write) is overkill here; we only need to read the file, so stick withr. - Printing a string instead of the variable: Your final line prints the literal string
"words"instead of the actual list of words you're trying to collect. - No dedicated unique word storage: You never create a separate list to hold all unique words from the entire file—you're reusing the line-specific
wordslist, which gets overwritten each loop iteration.
Corrected Working Code
Here's the fixed version that hits all the task requirements:
# Start with an empty list to store all unique words unique_words = [] # Open the file in read-only mode (using `with` ensures it gets closed automatically) with open("romeo.txt", "r") as fhand: # Loop through each line in the file for line in fhand: # Split the line into individual words (split() handles whitespace by default) line_words = line.split() # Check each word in the current line for word in line_words: # Only add the word if it's not already in our unique list if word not in unique_words: unique_words.append(word) # Sort the unique words alphabetically unique_words.sort() # Print the final sorted list print(unique_words)
Why This Works Better
withstatement: This is the Pythonic way to handle files—it automatically closes the file once we're done processing, so we don't have to remember to callfhand.close().- Dedicated
unique_wordslist: This list persists across all line iterations, so we're tracking unique words for the entire file, not just one line at a time. - Correct duplicate check: We check against
unique_words(our global list) instead of the line-specificline_wordslist, so we don't add duplicates. - Proper output: We print the actual
unique_wordsvariable, not a hardcoded string.
Bonus: Faster Deduplication with Sets
If you ever work with larger files (like full Shakespearean texts), checking if word not in unique_words can be slow because lists have linear lookup times. A set is designed for fast uniqueness checks—here's a more efficient version using a set:
# Sets automatically handle uniqueness, so we don't need manual checks unique_words = set() with open("romeo.txt", "r") as fhand: for line in fhand: for word in line.split(): unique_words.add(word) # Convert the set to a sorted list (sets aren't ordered) sorted_words = sorted(unique_words) print(sorted_words)
This does the same job but runs faster for large datasets—win-win!
内容的提问来源于stack exchange,提问作者Ssenkinga Abdul
相关产品推荐
相关产品推荐

