Python计算年龄标准化率时出现NaN值问题求助
解决年龄标准化率计算中出现NaN的问题
我在求职项目中处理疾病相关数据,前期计算乌干达和美国的分年龄人口权重百分比时结果正常,无NaN值。但后续用COPD表中的死亡率乘以对应权重计算年龄标准化率时,部分行出现NaN值。已排除除零情况,仅新计算的列存在该问题。
相关代码与数据
无错误的权重计算代码
# Calculation of the weight of the population by age #Uganda # Sum of observations from the column "Estimates, Uganda, 2019" total_observation = transposed_df["Estimates, Uganda, 2019"].sum() # The weight in percentage for each observation transposed_df["Weight Percentage Uganda"] = (transposed_df["Estimates, Uganda, 2019"] / total_observation) ##### #United States of America # Sum of observations from the column "Estimates, United States of America, 2019" total_observation1 = transposed_df["Estimates, United States of America, 2019"].sum() # The weight in percentage for each observation transposed_df["Weight Percentage United States of America"] = (transposed_df["Estimates, United States of America, 2019"] / total_observation1) #print transposed_df
出现NaN的年龄标准化率计算代码
#Age-standardized rate from Uganda transposed_df["Age-standardized rate from Uganda"] = (COPD["Death rate, Uganda, 2019"] * transposed_df["Weight Percentage Uganda"]) #Age-standardized rate from United States of America transposed_df["Age-standardized rate from United States of America"] = (COPD["Death rate, United States, 2019"] * transposed_df["Weight Percentage United States of America"]) #display result transposed_df
COPD数据表
| Age group (years) | Death rate, United States, 2019 | Death rate, Uganda, 2019 |
|---|---|---|
| 0-4 | 0.04 | 0.40 |
| 5-9 | 0.02 | 0.17 |
| 10-14 | 0.02 | 0.07 |
| 15-19 | 0.02 | 0.23 |
| 20-24 | 0.06 | 0.38 |
| 25-29 | 0.11 | 0.40 |
| 30-34 | 0.29 | 0.75 |
| 35-39 | 0.56 | 1.11 |
| 40-44 | 1.42 | 2.04 |
| 45-49 | 4.00 | 5.51 |
| 50-54 | 14.13 | 13.26 |
| 55-59 | 37.22 | 33.25 |
| 60-64 | 66.48 | 69.62 |
| 65-69 | 108.66 | 120.78 |
| 70-74 | 213.10 | 229.88 |
| 75-79 | 333.06 | 341.06 |
| 80-84 | 491.10 | 529.31 |
| 85+ | 894.45 | 710.40 |
排查情况:仅新计算的年龄标准化率列出现NaN,前期权重计算正常,已排除除零问题。
可能的原因及解决方案
1. 索引不匹配(最常见原因)
如果COPD和transposed_df的行索引未基于相同年龄组,或索引顺序不一致,相乘时会因无法匹配对应行产生NaN。
解决方法:统一两个DataFrame的索引为年龄组,确保对齐:
# 将COPD的索引设置为年龄组 COPD.set_index("Age group (years)", inplace=True) # 确保transposed_df的索引也是年龄组(若未设置) transposed_df.set_index("Age group (years)", inplace=True) # 重新计算年龄标准化率 transposed_df["Age-standardized rate from Uganda"] = COPD["Death rate, Uganda, 2019"] * transposed_df["Weight Percentage Uganda"] transposed_df["Age-standardized rate from United States of America"] = COPD["Death rate, United States, 2019"] * transposed_df["Weight Percentage United States of America"]
2. COPD表存在隐藏NaN值
展示的数据表无NaN,但实际加载的COPD数据可能存在缺失值。先检查:
# 检查COPD表中的NaN值 print(COPD.isna().sum())
若发现NaN,可根据需求填充:
# 用0填充NaN COPD.fillna(0, inplace=True) # 或用列均值填充 COPD.fillna(COPD.mean(), inplace=True)
3. 数据类型不兼容
若死亡率列为字符串类型,相乘时会自动转为NaN。检查并转换数据类型:
# 查看COPD表的数据类型 print(COPD.dtypes) # 转换为数值型 COPD["Death rate, United States, 2019"] = pd.to_numeric(COPD["Death rate, United States, 2019"], errors="coerce") COPD["Death rate, Uganda, 2019"] = pd.to_numeric(COPD["Death rate, Uganda, 2019"], errors="coerce")
4. 手动对齐行顺序
若索引设置后仍有问题,可按年龄组排序确保行顺序一致:
# 按年龄组排序 COPD.sort_index(inplace=True) transposed_df.sort_index(inplace=True)
内容的提问来源于stack exchange,提问作者Giovanna
相关产品推荐
相关产品推荐

