You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效遍历列表统计医保数据各区域吸烟人群占比?

优化医保数据集吸烟占比统计的高效方案

你的代码确实存在冗余问题——硬编码大量变量、重复判断逻辑,还容易出现笔误(比如最后一个东南非吸烟者的比例计算错用了西北的数值)。针对你的需求,这里提供几个逐步进阶的优化方案,适合刚学Python两周的你上手:

方案一:用嵌套字典替代硬编码变量

通过字典动态存储不同区域的吸烟/非吸烟人数,彻底避免重复定义变量:

from collections import defaultdict

def smoker_region_diff(smoker_status, region):
    # 初始化嵌套字典:区域 -> 吸烟状态 -> 人数
    count_dict = defaultdict(lambda: defaultdict(int))
    total = len(smoker_status)
    
    # 遍历数据统计
    for status, reg in zip(smoker_status, region):
        count_dict[reg][status] += 1
    
    # 循环输出各区域占比
    region_names = {
        'northwest': '西北',
        'northeast': '东北',
        'southwest': '西南',
        'southeast': '东南'
    }
    for reg, cn_name in region_names.items():
        smokers = count_dict[reg]['yes']
        non_smokers = count_dict[reg]['no']
        print(f'{cn_name}区域吸烟者占比: {smokers/total:.2%}')
        print(f'{cn_name}区域非吸烟者占比: {non_smokers/total:.2%}')

# 直接传入两个列表即可调用
smoker_region_diff(smoker_status, region)

方案二:用collections.Counter直接统计组合

Counter工具可以一键统计(吸烟状态, 区域)的出现次数,代码更简洁:

from collections import Counter

def smoker_region_diff(smoker_status, region):
    total = len(smoker_status)
    # 统计所有(状态, 区域)组合的次数
    status_region_counter = Counter(zip(smoker_status, region))
    
    region_names = {
        'northwest': '西北',
        'northeast': '东北',
        'southwest': '西南',
        'southeast': '东南'
    }
    status_labels = {'yes': '吸烟者', 'no': '非吸烟者'}
    
    # 遍历输出
    for reg, cn_name in region_names.items():
        for status, label in status_labels.items():
            count = status_region_counter.get((status, reg), 0)
            print(f'{cn_name}区域{label}占比: {count/total:.2%}')

smoker_region_diff(smoker_status, region)

方案三:用Pandas快速分析(推荐处理数据集)

如果你的医保数据是表格形式,Pandas是最适合的工具,几行代码就能完成专业级统计:

import pandas as pd

# 把列表转换成DataFrame(表格结构)
df = pd.DataFrame({
    'smoker': smoker_status,
    'region': region
})

# 计算各区域-状态组合相对于总人数的占比
proportion = df.groupby(['region', 'smoker']).size() / len(df)

# 格式化输出
region_names = {
    'northwest': '西北',
    'northeast': '东北',
    'southwest': '西南',
    'southeast': '东南'
}
status_labels = {'yes': '吸烟者', 'no': '非吸烟者'}

for (reg, status), prop in proportion.items():
    print(f'{region_names[reg]}区域{status_labels[status]}占比: {prop:.2%}')

额外提示

如果需要计算区域内部的吸烟占比(而非相对于总人数),可以用Pandas的normalize=True参数:

region_inner_proportion = df.groupby('region')['smoker'].value_counts(normalize=True)

内容的提问来源于stack exchange,提问作者whorrodwi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 22:22:45