You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定索引合并DataFrame行并排除NaN值?

如何将含NaN的行与前一行合并(排除NaN值)

问题描述

现有如下数据集(实际包含更多列和行):

id  type      value
0  104     0       7999
1  105     1  196193579
2  108     0     245744
3  NaN     1        NaN

部分行存在NaN值,已获取这些行的索引(例如indexes=[3]),需要将这些行与前一行合并,合并时排除NaN值,最终得到的DataFrame如下:

id type      value
0  104    0       7999
1  105    1  196193579
2  108   01     245744

注意:首行绝不会出现在索引列表中,解决方案需基于给定的索引列表,也可使用已知的含NaN的列名。

解决方案

实现思路

  1. 倒序遍历目标索引:避免删除行后索引偏移导致后续处理出错
  2. 逐列处理:对每个目标行,将其非NaN值与前一行的对应列值拼接(或替换前一行的NaN值)
  3. 删除处理后的目标行,最后重置索引(可选)

代码实现

基础版本

import pandas as pd
import numpy as np

# 示例数据集
df = pd.DataFrame({
    'id': [104, 105, 108, np.nan],
    'type': [0, 1, 0, 1],
    'value': [7999, 196193579, 245744, np.nan]
})

def merge_nan_rows(df, indexes):
    # 倒序遍历索引,防止删除行后索引混乱
    for idx in sorted(indexes, reverse=True):
        prev_idx = idx - 1
        # 遍历所有列
        for col in df.columns:
            curr_val = df.loc[idx, col]
            prev_val = df.loc[prev_idx, col]
            
            # 仅保留非NaN值,若两行都有值则拼接
            if not pd.isna(curr_val):
                if not pd.isna(prev_val):
                    df.loc[prev_idx, col] = str(prev_val) + str(curr_val)
                else:
                    df.loc[prev_idx, col] = curr_val
        # 删除当前行
        df = df.drop(idx)
    
    # 重置索引(按需选择)
    df = df.reset_index(drop=True)
    return df

# 调用示例
indexes = [3]
result_df = merge_nan_rows(df.copy(), indexes)
print(result_df)

优化版本(指定含NaN的列)

如果已知存在NaN的列名,可以只处理这些列,提升效率:

def merge_nan_rows(df, indexes, nan_cols=None):
    # 只处理指定的含NaN列,默认处理所有列
    cols_to_process = nan_cols if nan_cols is not None else df.columns
    
    for idx in sorted(indexes, reverse=True):
        prev_idx = idx - 1
        for col in cols_to_process:
            curr_val = df.loc[idx, col]
            prev_val = df.loc[prev_idx, col]
            
            if not pd.isna(curr_val):
                if not pd.isna(prev_val):
                    df.loc[prev_idx, col] = str(prev_val) + str(curr_val)
                else:
                    df.loc[prev_idx, col] = curr_val
        df = df.drop(idx)
    
    df = df.reset_index(drop=True)
    return df

# 调用示例(已知含NaN的列是id和value)
indexes = [3]
result_df = merge_nan_rows(df.copy(), indexes, nan_cols=['id', 'value'])

输出结果

运行上述代码后,得到的结果符合预期:

id type      value
0  104    0       7999
1  105    1  196193579
2  108   01     245744

内容的提问来源于stack exchange,提问作者Learning from masters

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:56:25