基于索引值对DataFrame全列执行乘法运算的实现方法
问题描述
有两个均以Age为索引的DataFrame,需要将倍率数据中对应年龄的倍率,应用到人口数据的各城镇列上。
数据示例
倍率数据(单列DataFrame)
import pandas as pd import numpy as np rate_for_yr = pd.DataFrame(np.array([[15,90],[16,80],[17,70]]), columns=["Age","2020"]) rate_for_yr = rate_for_yr.set_index('Age')
| Age | 2020 |
|---|---|
| 15 | 90 |
| 16 | 80 |
人口数据(多列DataFrame)
pop_by_town = pd.DataFrame(np.array([[15, 1, 2, 3, 4, 5, 6], [16, 7, 8, 9, 10, 11, 12], [17, 13, 14, 15, 16, 17, 18]]), columns=["Age","townA", "townB", "townC", "townD", "townE", "townF"]) pop_by_town = pop_by_town.set_index('Age')
| Age | townA | townB | townC |
|---|---|---|---|
| 15 | 1 | 2 | 3 |
| 16 | 7 | 8 | 9 |
期望输出
| Age | townA | townB | townC |
|---|---|---|---|
| 15 | 90 | 180 | 270 |
| 16 | 560 | 640 | 720 |
解决方案
利用pandas的索引对齐与广播特性,直接将人口DataFrame与倍率Series相乘即可,自动处理索引匹配与多列批量计算:
# 提取倍率列作为Series,与人口DataFrame按索引相乘 result = pop_by_town * rate_for_yr['2020'] # 可选:过滤出需要的城镇列(如示例中的前3列) result = result[['townA', 'townB', 'townC']] # 可选:移除索引不匹配导致的NaN行(如Age=17) result = result.dropna() print(result)
原理说明
- 两个对象共享
Age索引,pandas会自动匹配相同索引的行进行计算 - 单列倍率Series会被广播到人口DataFrame的所有列,实现逐行对应的倍率批量应用
- 索引不匹配的行(如Age=17)会生成NaN,可通过
dropna()过滤
内容的提问来源于stack exchange,提问作者GlassShark1
相关产品推荐
相关产品推荐

