如何在Pandas的lambda函数中访问上一行值或df.apply()中获取行索引
问题1:生成Charge和VisitorsCharge列
需求描述
现有如下Pandas DataFrame:
import pandas as pd df = pd.DataFrame({'Persons':[10,20,30], 'Bill':[110,240,365], 'Guests':[12,25,29],'Visitors':[15,23,27]})
输出结构:
| Persons | Bill | Guests | Visitors |
|---|---|---|---|
| 10 | 110 | 12 | 15 |
| 20 | 240 | 25 | 23 |
| 30 | 365 | 29 | 27 |
需要生成包含Charge和VisitorsCharge列的新DataFrame:
| Persons | Bill | Guests | Visitors | Charge | VisitorsCharge |
|---|---|---|---|---|---|
| 10 | 110 | 12 | 15 | 136 | 175 |
| 20 | 240 | 25 | 23 | 302.5 | 277.5 |
| 30 | 365 | 29 | 27 | 352.5 | 327.5 |
计算规则:
Charge:以Persons和Bill为参考,对Guests做线性插值- 非最后一行:用当前行和下一行的
Persons、Bill数据,通过scipy.stats.linregress计算斜率与截距,代入Guests得到结果 - 最后一行:用当前行和上一行的数据计算
- 非最后一行:用当前行和下一行的
VisitorsCharge:逻辑与Charge一致,仅将Guests替换为Visitors
解决方案
通过遍历行索引直接定位相邻行数据,结合线性回归计算:
import pandas as pd from scipy.stats import linregress df = pd.DataFrame({'Persons':[10,20,30], 'Bill':[110,240,365], 'Guests':[12,25,29],'Visitors':[15,23,27]}) def calculate_charge(row_idx, df, target_col): if row_idx < len(df)-1: # 非最后一行:取当前行和下一行数据 x = df.loc[[row_idx, row_idx+1], 'Persons'].values y = df.loc[[row_idx, row_idx+1], 'Bill'].values else: # 最后一行:取当前行和上一行数据 x = df.loc[[row_idx-1, row_idx], 'Persons'].values y = df.loc[[row_idx-1, row_idx], 'Bill'].values # 计算线性回归参数 slope, intercept, _, _, _ = linregress(x, y) # 代入目标列值计算结果 return slope * df.loc[row_idx, target_col] + intercept # 生成Charge列 df['Charge'] = [calculate_charge(i, df, 'Guests') for i in df.index] # 生成VisitorsCharge列 df['VisitorsCharge'] = [calculate_charge(i, df, 'Visitors') for i in df.index] print(df)
问题2:生成列C
需求描述
现有如下DataFrame:
| A | B |
|---|---|
| 1 | 100 |
| 2 | 200 |
| 3 | 300 |
需要生成包含列C的DataFrame:
| A | B | C |
|---|---|---|
| 1 | 100 | '1-2-100-200' |
| 2 | 200 | '2-3-200-300' |
| 3 | 300 | '2-3-200-300' |
解决方案
通过shift生成相邻行数据后拼接字符串,实现更简洁高效:
import pandas as pd df = pd.DataFrame({'A':[1,2,3], 'B':[100,200,300]}) # 生成下一行的A、B值,最后一行填充为倒数第二行的对应值 df['next_A'] = df['A'].shift(-1).fillna(df['A'].iloc[-2]) df['next_B'] = df['B'].shift(-1).fillna(df['B'].iloc[-2]) # 拼接生成列C df['C'] = df.apply(lambda row: f"{row['A']}-{row['next_A']}-{row['B']}-{row['next_B']}", axis=1) # 移除辅助列(可选) df.drop(['next_A', 'next_B'], axis=1, inplace=True) print(df)
也可以直接通过索引判断处理:
df['C'] = df.apply(lambda row: f"{row['A']}-{df.loc[row.name+1, 'A']}-{row['B']}-{df.loc[row.name+1, 'B']}" if row.name < len(df)-1 else df.loc[row.name-1, 'C'], axis=1)
内容的提问来源于stack exchange,提问作者moys
相关产品推荐
相关产品推荐

