如何用Python从CSV文件识别并统计唯一国家代码的出现次数
嘿,这个需求很常见,我给你分享几种实用的实现方式,从基础手动统计到高效工具都有~
方法1:手动用字典统计(基础版)
你可以自己初始化一个空字典,遍历过程中逐个更新每个国家代码的计数,适合理解底层逻辑:
import csv country_counts = {} with open('2017CountryData.csv') as csvfile: readCSV = csv.reader(csvfile, delimiter=',') # 重要提醒:如果你的CSV第一行是表头(比如"Country Code"这类),记得先跳过这行 next(readCSV) # 去掉这行如果没有表头 for row in readCSV: country_code = row[1] # 判断代码是否已在字典中,更新计数 if country_code in country_counts: country_counts[country_code] += 1 else: country_counts[country_code] = 1 # 打印最终统计结果 for code, count in country_counts.items(): print(f"{code}: {count}")
方法2:用collections.Counter(推荐,简洁高效)
Python标准库的collections.Counter就是专门用来做元素计数的工具,代码会简洁很多:
import csv from collections import Counter with open('2017CountryData.csv') as csvfile: readCSV = csv.reader(csvfile, delimiter=',') next(readCSV) # 跳过表头(按需保留) # 用生成器表达式直接把所有国家代码传给Counter country_counts = Counter(row[1] for row in readCSV) # 打印所有统计结果 for code, count in country_counts.items(): print(f"{code}: {count}") # 额外技巧:如果想找出现次数最多的前N个国家代码,用most_common() # print(country_counts.most_common(3)) # 输出出现次数前三的代码和计数
方法3:用pandas(适合大数据量场景)
如果你的CSV文件数据量很大,或者平时习惯用pandas处理数据,这个方法会更高效,代码也极简:
import pandas as pd # 读取CSV文件,统计第二列(索引为1)的元素出现次数 country_counts = pd.read_csv('2017CountryData.csv').iloc[:, 1].value_counts() # 打印结果,pandas会自动按计数从高到低排序 print(country_counts)
内容的提问来源于stack exchange,提问作者pythonlearner1994
相关产品推荐
相关产品推荐

