如何高效遍历列表统计医保数据各区域吸烟人群占比?
优化医保数据集吸烟占比统计的高效方案
你的代码确实存在冗余问题——硬编码大量变量、重复判断逻辑,还容易出现笔误(比如最后一个东南非吸烟者的比例计算错用了西北的数值)。针对你的需求,这里提供几个逐步进阶的优化方案,适合刚学Python两周的你上手:
方案一:用嵌套字典替代硬编码变量
通过字典动态存储不同区域的吸烟/非吸烟人数,彻底避免重复定义变量:
from collections import defaultdict def smoker_region_diff(smoker_status, region): # 初始化嵌套字典:区域 -> 吸烟状态 -> 人数 count_dict = defaultdict(lambda: defaultdict(int)) total = len(smoker_status) # 遍历数据统计 for status, reg in zip(smoker_status, region): count_dict[reg][status] += 1 # 循环输出各区域占比 region_names = { 'northwest': '西北', 'northeast': '东北', 'southwest': '西南', 'southeast': '东南' } for reg, cn_name in region_names.items(): smokers = count_dict[reg]['yes'] non_smokers = count_dict[reg]['no'] print(f'{cn_name}区域吸烟者占比: {smokers/total:.2%}') print(f'{cn_name}区域非吸烟者占比: {non_smokers/total:.2%}') # 直接传入两个列表即可调用 smoker_region_diff(smoker_status, region)
方案二:用collections.Counter直接统计组合
Counter工具可以一键统计(吸烟状态, 区域)的出现次数,代码更简洁:
from collections import Counter def smoker_region_diff(smoker_status, region): total = len(smoker_status) # 统计所有(状态, 区域)组合的次数 status_region_counter = Counter(zip(smoker_status, region)) region_names = { 'northwest': '西北', 'northeast': '东北', 'southwest': '西南', 'southeast': '东南' } status_labels = {'yes': '吸烟者', 'no': '非吸烟者'} # 遍历输出 for reg, cn_name in region_names.items(): for status, label in status_labels.items(): count = status_region_counter.get((status, reg), 0) print(f'{cn_name}区域{label}占比: {count/total:.2%}') smoker_region_diff(smoker_status, region)
方案三:用Pandas快速分析(推荐处理数据集)
如果你的医保数据是表格形式,Pandas是最适合的工具,几行代码就能完成专业级统计:
import pandas as pd # 把列表转换成DataFrame(表格结构) df = pd.DataFrame({ 'smoker': smoker_status, 'region': region }) # 计算各区域-状态组合相对于总人数的占比 proportion = df.groupby(['region', 'smoker']).size() / len(df) # 格式化输出 region_names = { 'northwest': '西北', 'northeast': '东北', 'southwest': '西南', 'southeast': '东南' } status_labels = {'yes': '吸烟者', 'no': '非吸烟者'} for (reg, status), prop in proportion.items(): print(f'{region_names[reg]}区域{status_labels[status]}占比: {prop:.2%}')
额外提示
如果需要计算区域内部的吸烟占比(而非相对于总人数),可以用Pandas的normalize=True参数:
region_inner_proportion = df.groupby('region')['smoker'].value_counts(normalize=True)
内容的提问来源于stack exchange,提问作者whorrodwi
相关产品推荐
相关产品推荐

