如何在Python中统计列表元素出现频率并跳过空值?
报错原因
- 实际运行代码中循环变量和计数变量重名:你先定义了计数变量
c = 0,但实际循环写的是for c in countries,这会让c在循环过程中被赋值为列表里的国家字符串,后续执行c += 1时就会出现「字符串不能和整数相加」的类型错误。 - 现有代码逻辑也不符合统计频次的需求,仅做了遍历计数,没有实现去重统计、也没有过滤空值的逻辑。
正确实现方案
方法1:用collections.Counter最简实现
from collections import Counter import pandas as pd # 若已导入过pandas可跳过该行 # 一步合并两列数据,无需循环append countries = pd.concat([survey['C1'], survey['C2']]).tolist() # 过滤空值:同时兼容pandas原生NaN和字符串类型的'nan' filtered_countries = [i for i in countries if pd.notna(i) and str(i).strip() != 'nan'] # 统计频次 count_result = Counter(filtered_countries) # 打印结果 for country, cnt in count_result.items(): print(f"{country}: {cnt}")
方法2:手动实现统计逻辑
# 合并两列数据 countries = [] for val in survey['C1']: countries.append(val) for val in survey['C2']: countries.append(val) count_dict = {} for item in countries: # 跳过空值 if pd.isna(item) or str(item).strip() == 'nan': continue # 频次统计逻辑 if item in count_dict: count_dict[item] += 1 else: count_dict[item] = 1 # 打印结果 for country, cnt in count_dict.items(): print(f"{country} 出现次数:{cnt}")
注意事项
如果读入Excel时空值被统一识别为字符串类型的'nan',直接用if item == 'nan'判断跳过即可;如果是pandas原生的NaN缺失值,用pd.isna(item)判断准确性更高。
内容的提问来源于stack exchange,提问作者Natalia Resende
相关产品推荐
相关产品推荐

