7.8实验:词频统计(列表与CSV)——如何读取CSV并去重?
问题分析与修复方案
原代码的核心问题
- 重复输出同一单词:遍历每个单词时都会打印一次统计结果,导致同一个单词多次重复出现(比如
cat出现2次,代码会输出两次cat 2)。 - 统计效率低下:每次调用
lines.count(w)都会重新遍历整行列表,做了大量重复计算。 - 未支持多行CSV:如果文件有多行,原代码只会逐行统计,不会合并所有行的单词做全局统计。
修复后的代码实现
方案1:严格区分大小写统计
用字典来记录每个单词的出现次数,遍历所有单词时更新计数,最后统一输出无重复的结果:
import csv user_input = input() word_counts = {} with open(user_input, 'r') as name_CSV: paper_copy = csv.reader(name_CSV) for line in paper_copy: for word in line: # 清理单词前后可能的空白字符 cleaned_word = word.strip() if cleaned_word: # 更新计数:存在则加1,不存在则初始化为1 word_counts[cleaned_word] = word_counts.get(cleaned_word, 0) + 1 # 遍历字典输出无重复的单词和频次 for word, count in word_counts.items(): print(word, count)
方案2:忽略大小写统计(Hello与hello视为同一单词)
如果需要忽略大小写差异,只需在统计前将单词统一转为小写(或大写):
import csv user_input = input() word_counts = {} with open(user_input, 'r') as name_CSV: paper_copy = csv.reader(name_CSV) for line in paper_copy: for word in line: cleaned_word = word.strip().lower() if cleaned_word: word_counts[cleaned_word] = word_counts.get(cleaned_word, 0) + 1 for word, count in word_counts.items(): print(word, count)
测试效果
针对输入文件input1.csv的内容:
hello,cat,man,hey,dog,boy,Hello,man,cat,woman,dog,Cat,hey,boy
- 方案1(区分大小写)输出:
hello 1 cat 2 man 2 hey 2 dog 2 boy 2 Hello 1 woman 1 Cat 1
- 方案2(不区分大小写)输出:
hello 2 cat 3 man 2 hey 2 dog 2 boy 2 woman 1
内容的提问来源于stack exchange,提问作者KnowsNothing
相关产品推荐
相关产品推荐

