Python处理《爱丽丝梦游仙境》文本统计唯一单词数的优化问题
优化思路
- 直接用Python内置的
set(集合)存储处理后的单词,集合天然支持元素去重,最终集合的长度就是唯一单词的数量,统计效率极高 - 移除多余的
split()操作:你当前遍历的wordlist每个元素本身就是单个单词,处理后已经是独立的单词字符串,不需要再做拆分 - 增加空值过滤:避免全标点的字符串处理后变成空值被误统计
原代码问题说明
你原来的写法中,每次循环都会生成一个仅包含当前单词的单元素列表newlist,没有做内容累积,所以无法得到统一的单词集合;同时split()操作对于单个单词的字符串来说完全是冗余操作,没有实际作用。
优化后代码
from string import punctuation def count_unique_words(wordlist): # 初始化空集合用于存储去重后的单词 unique_words = set() for word in wordlist: # 去除两端标点+转小写 processed_word = word.strip(punctuation).lower() # 过滤空字符串后加入集合 if processed_word: unique_words.add(processed_word) # 集合长度就是唯一单词数 return len(unique_words) # 调用示例,直接传入你的wordlist即可 print(count_unique_words(wordlist))
如果追求更简洁的写法,可以用集合推导式实现:
from string import punctuation unique_count = len({word.strip(punctuation).lower() for word in wordlist if word.strip(punctuation).lower()})
直接读取txt文件的高效实现
如果你不想提前生成wordlist,可以直接读取文件内容处理,内存占用更低:
from string import punctuation def count_unique_words_from_file(file_path): unique_words = set() with open(file_path, 'r', encoding='utf-8') as f: for line in f: # 按空格拆分每行的单词 for word in line.split(): processed_word = word.strip(punctuation).lower() if processed_word: unique_words.add(processed_word) return len(unique_words) # 调用示例,填入你的txt文件路径 print(count_unique_words_from_file("Alice in the Wonderland.txt"))
内容的提问来源于stack exchange,提问作者gavmross
相关产品推荐
相关产品推荐

