You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

7.8实验:词频统计(列表与CSV)——如何读取CSV并去重?

问题分析与修复方案

原代码的核心问题

  • 重复输出同一单词:遍历每个单词时都会打印一次统计结果,导致同一个单词多次重复出现(比如cat出现2次,代码会输出两次cat 2)。
  • 统计效率低下:每次调用lines.count(w)都会重新遍历整行列表,做了大量重复计算。
  • 未支持多行CSV:如果文件有多行,原代码只会逐行统计,不会合并所有行的单词做全局统计。

修复后的代码实现

方案1:严格区分大小写统计

用字典来记录每个单词的出现次数,遍历所有单词时更新计数,最后统一输出无重复的结果:

import csv

user_input = input()
word_counts = {}

with open(user_input, 'r') as name_CSV:
    paper_copy = csv.reader(name_CSV)
    for line in paper_copy:
        for word in line:
            # 清理单词前后可能的空白字符
            cleaned_word = word.strip()
            if cleaned_word:
                # 更新计数:存在则加1,不存在则初始化为1
                word_counts[cleaned_word] = word_counts.get(cleaned_word, 0) + 1

# 遍历字典输出无重复的单词和频次
for word, count in word_counts.items():
    print(word, count)

方案2:忽略大小写统计(Hello与hello视为同一单词)

如果需要忽略大小写差异,只需在统计前将单词统一转为小写(或大写):

import csv

user_input = input()
word_counts = {}

with open(user_input, 'r') as name_CSV:
    paper_copy = csv.reader(name_CSV)
    for line in paper_copy:
        for word in line:
            cleaned_word = word.strip().lower()
            if cleaned_word:
                word_counts[cleaned_word] = word_counts.get(cleaned_word, 0) + 1

for word, count in word_counts.items():
    print(word, count)

测试效果

针对输入文件input1.csv的内容:

hello,cat,man,hey,dog,boy,Hello,man,cat,woman,dog,Cat,hey,boy

  • 方案1(区分大小写)输出:
hello 1
cat 2
man 2
hey 2
dog 2
boy 2
Hello 1
woman 1
Cat 1
  • 方案2(不区分大小写)输出:
hello 2
cat 3
man 2
hey 2
dog 2
boy 2
woman 1

内容的提问来源于stack exchange,提问作者KnowsNothing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:18:19