You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中删除NaN值并将缺失行数据累加到下一行非空行

问题解决:合并含NaN行的累加值并保留有效行

原始数据

index   date    B   N   S   y_B   y_N   y_S  f_price    y_price  p_change   price_day_t-2
1   2021-01-12  1   29  0   2.0   57.0  0.0  10250.0    10760.0  -1.0        11060.0
2   2021-01-13  0   67  0   1.0   29.0  0.0  9810.0     10250.0  -1.0        10760.0
3   2021-01-14  2   19  0   0.0   67.0  0.0  NaN         NaN     NaN           NaN
4   2021-01-15  1   6   0   2.0   19.0  0.0  NaN         NaN     NaN           NaN
5   2021-01-16  2   46  0   1.0   6.0   0.0  9340.0     9810.0   -1.0        10250.0
6   2021-01-17  3   22  0   2.0   46.0  0.0  NaN         NaN     NaN           NaN
7   2021-01-18  1   34  0   3.0   22.0  0.0  8890.0     9340.0   -1.0        9810.0

需求说明

删除所有包含NaN值的行,同时将这些NaN行的B、N、S、y_B、y_N、y_S列的值累加至下一个不含NaN的行中,其他列(如date、f_price等)保留非NaN行的原值,最终得到如下结果:

index   date    B   N   S   y_B   y_N   y_S  f_price    y_price  p_change   price_day_t-2
1   2021-01-12  1   29  0   2.0   57.0  0.0  10250.0    10760.0  -1.0        11060.0
2   2021-01-13  0   67  0   1.0   29.0  0.0  9810.0     10250.0  -1.0        10760.0
5   2021-01-16  5   71  0   3.0   92.0  0.0  9340.0     9810.0   -1.0        10250.0
7   2021-01-18  4   56  0   5.0   68.0  0.0  8890.0     9340.0   -1.0        9810.0

实现方案

通过分组累加+保留有效行属性的方式实现,具体步骤如下:

1. 导入依赖并构造DataFrame

确保已导入pandas,然后构造原始数据:

import pandas as pd

data = {
    'index': [1,2,3,4,5,6,7],
    'date': ['2021-01-12','2021-01-13','2021-01-14','2021-01-15','2021-01-16','2021-01-17','2021-01-18'],
    'B': [1,0,2,1,2,3,1],
    'N': [29,67,19,6,46,22,34],
    'S': [0,0,0,0,0,0,0],
    'y_B': [2.0,1.0,0.0,2.0,1.0,2.0,3.0],
    'y_N': [57.0,29.0,67.0,19.0,6.0,46.0,22.0],
    'y_S': [0.0,0.0,0.0,0.0,0.0,0.0,0.0],
    'f_price': [10250.0,9810.0,None,None,9340.0,None,8890.0],
    'y_price': [10760.0,10250.0,None,None,9810.0,None,9340.0],
    'p_change': [-1.0,-1.0,None,None,-1.0,None,-1.0],
    'price_day_t-2': [11060.0,10760.0,None,None,10250.0,None,9810.0]
}
df = pd.DataFrame(data).set_index('index')

2. 创建分组键

标记有效行(无NaN的行),生成分组键将连续的NaN行和下一个非NaN行归为同一组:

# 标记所有无NaN的行
valid_rows = df.notna().all(axis=1)
# 从后往前累加有效行标记,确保每组最后一行是有效行
group_key = valid_rows[::-1].cumsum()[::-1]

3. 分组聚合处理

对需要累加的列求和,对其他列取每组最后一个值(即有效行的原值):

# 定义需要累加的列
sum_cols = ['B', 'N', 'S', 'y_B', 'y_N', 'y_S']
# 定义需要保留最后一个值的列
last_cols = [col for col in df.columns if col not in sum_cols]

# 执行分组聚合
result = df.groupby(group_key).agg({
    **{col: 'sum' for col in sum_cols},
    **{col: 'last' for col in last_cols}
})

# 重置索引为原始有效行的index
result = result.set_index(df[valid_rows].index)

4. 查看结果

执行上述代码后,result即为目标DataFrame:

date  B   N  S  y_B  y_N  y_S  f_price  y_price  p_change  price_day_t-2
index                                                                               
1     2021-01-12  1  29  0  2.0  57.0  0.0  10250.0  10760.0      -1.0        11060.0
2     2021-01-13  0  67  0  1.0  29.0  0.0   9810.0  10250.0      -1.0        10760.0
5     2021-01-16  5  71  0  3.0  92.0  0.0   9340.0   9810.0      -1.0        10250.0
7     2021-01-18  4  56  0  5.0  68.0  0.0   8890.0   9340.0      -1.0         9810.0

内容的提问来源于stack exchange,提问作者Mohmmad Hadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 05:54:09