Python筛选DataFrame:保留特定死因关联数据以用于可视化
解决方法
1. 先明确数据结构
假设你的DataFrame是宽表格式,每一组年份对应两列:死因名称列和对应占比列,示例结构如下:
| 索引 | Causas 2 de año 2016 | % 1 | Causas 2 de año 2017 | % 2 | Causas 2 de año 2018 | % 3 |
|---|---|---|---|---|---|---|
| 0 | 心血管疾病 | 25% | 心血管疾病 | 26% | 心血管疾病 | 24% |
| 1 | 癌症 | 20% | 癌症 | 21% | 癌症 | 22% |
| 2 | 呼吸系统疾病 | 15% | 呼吸系统疾病 | 14% | 呼吸系统疾病 | 16% |
2. 提取特定死因的年度占比
方法一:手动匹配(适合年份少的场景)
假设你要提取心血管疾病的年度数据,直接针对每一年的列做筛选:
import pandas as pd # 假设你的原始数据存储在df变量中 target_cause = "心血管疾病" yearly_data = {} # 提取2016年数据 mask_2016 = df["Causas 2 de año 2016"] == target_cause yearly_data[2016] = df.loc[mask_2016, "% 1"].values[0] # 提取2017年数据 mask_2017 = df["Causas 2 de año 2017"] == target_cause yearly_data[2017] = df.loc[mask_2017, "% 2"].values[0] # 提取2018年数据(按此逻辑补充更多年份) mask_2018 = df["Causas 2 de año 2018"] == target_cause yearly_data[2018] = df.loc[mask_2018, "% 3"].values[0] # 转成适合绘图的长表格式 result_df = pd.DataFrame(list(yearly_data.items()), columns=["年份", "占比"]) # 把带%的字符串转成数值型 result_df["占比"] = result_df["占比"].str.replace("%", "").astype(float)
方法二:批量自动处理(适合年份多的场景)
如果年份较多,手动写代码效率低,可以通过列名规律自动配对死因列和占比列:
import pandas as pd target_cause = "心血管疾病" yearly_data = [] # 筛选所有死因列 cause_cols = [col for col in df.columns if "Causas 2 de año" in col] for cause_col in cause_cols: # 从列名提取年份 year = int(cause_col.split(" ")[-1]) # 获取对应占比列(假设占比列是死因列的下一列,可根据实际结构调整) col_idx = df.columns.get_loc(cause_col) percent_col = df.columns[col_idx + 1] # 筛选目标死因的占比 mask = df[cause_col] == target_cause percent_value = df.loc[mask, percent_col].values[0] yearly_data.append({"年份": year, "占比": percent_value}) # 转成DataFrame并处理数值 result_df = pd.DataFrame(yearly_data) result_df["占比"] = result_df["占比"].str.replace("%", "").astype(float)
3. 绘图示例
用Matplotlib绘制占比变化折线图:
import matplotlib.pyplot as plt plt.figure(figsize=(10, 6)) plt.plot(result_df["年份"], result_df["占比"], marker='o', linestyle='-', color='#1f77b4') plt.title(f"{target_cause} 年度占比变化") plt.xlabel("年份") plt.ylabel("占比 (%)") plt.grid(alpha=0.3) plt.show()
常见问题排查
- 如果
.loc筛选无结果,先检查布尔掩码:打印df["Causas 2 de año 2016"] == target_cause,确认是否有True值,可能是字符串大小写、空格或拼写不一致导致匹配失败。 - 若提取
values[0]报错,说明未找到对应死因,核对目标死因的拼写是否和DataFrame中的完全一致。
内容的提问来源于stack exchange,提问作者Luciano Carvajal
相关产品推荐
相关产品推荐

