Python列表与split方法使用问题:无法生成单一单词列表求助
Hey there! Let's break down why your code is outputting multiple sublists instead of a single, sorted list of unique words, and fix it step by step.
What's Wrong with the Original Code?
Looking at your code, the key issue is how you're handling the split words:
for line in fh: for each in line: # This loop is totally unnecessary! word = line.split() # This gives a LIST of words for the entire line if word not in lst: lst.append(word) # You're adding the whole list as a single element
Instead of adding individual words to your main list lst, you're adding entire lists of words from each line—that's why you end up with multiple separate lists in the final output. The inner for each in line loop is also redundant, since it iterates over every character in the line (which you don't need at all).
Solution 1: Manual Duplicate Check with a List
Here's a corrected version that adds individual words and ensures no duplicates, then sorts the final list:
fname = input("Enter file name: ") fh = open(fname) lst = [] for line in fh: # Strip extra whitespace (like newlines) and split into individual words words = line.strip().split() # Iterate over each word in the line's word list for word in words: # Only add the word if it's not already in the list if word not in lst: lst.append(word) # Sort the list alphabetically before printing lst.sort() print(lst)
Solution 2: Use a Set for Automatic Deduplication (More Efficient)
For a cleaner and faster approach, use a Python set—sets automatically ignore duplicate values, so you don't need to manually check if a word exists:
fname = input("Enter file name: ") # Use `with` statement to auto-close the file (safer practice!) with open(fname) as fh: word_set = set() for line in fh: words = line.strip().split() # Add all words from the line to the set (duplicates are ignored) word_set.update(words) # Convert the set to a sorted list sorted_words = sorted(word_set) print(sorted_words)
This method is more efficient, especially for large files, because checking membership in a set is much faster than in a list.
Key Takeaways
- Always iterate over individual words (not the entire list of words from a line) when building your unique word collection.
- Use
strip()before splitting to avoid accidental empty strings from leading/trailing whitespace. - For deduplication, sets are your friend—they handle the heavy lifting for you!
内容的提问来源于stack exchange,提问作者Shivam Rastogi

