You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas分组内对多列分别采用向前填充与插值补全缺失值?

Groupby Pandas DataFrame with Column-Specific Missing Value Imputation

Hey there, great question! Handling column-specific imputation within groups is a common scenario, and luckily Pandas gives us a flexible way to tackle this. Let's break down the solution using your sample data first, then scale it to your full 12-column dataset.

Step 1: Sample Data Setup

First, let's recreate your sample DataFrame to work with:

import pandas as pd
import numpy as np

data = {
    'brake': [0.0, 1.0, np.nan, np.nan, 0.0, np.nan, np.nan, 1.0, 0.0, np.nan, 1.0, 0.0, np.nan, 1.0, np.nan, 0.0, np.nan, np.nan, np.nan, np.nan],
    'speed': [np.nan, np.nan, 1.264, 0.000, np.nan, 1.264, 6.704, np.nan, np.nan, 11.746, np.nan, np.nan, 16.961, np.nan, 11.832, np.nan, 17.082, 22.435, 28.707, 34.216],
    'trip': [1,1,1,1,1,1,1,1,1,2,2,2,2,3,3,3,3,3,3,3]
}

df = pd.DataFrame(data)

Step 2: Define Column-Specific Imputation Rules

Create a dictionary where each key is a column name, and the value is the imputation logic you want to apply to that column within each group. For your requirements, plus examples for other columns you might have:

# Define your custom imputation strategy per column
impute_strategy = {
    # Your required rules
    'brake': lambda x: x.ffill(),  # Forward fill (carry last known value)
    'speed': lambda x: x.interpolate(method='linear'),  # Linear interpolation
    
    # Add rules for your remaining 10 columns here:
    # Example 1: Backward fill another column
    # 'steering': lambda x: x.bfill(),
    # Example 2: Fill with column mean per group
    # 'distance': lambda x: x.fillna(x.mean()),
    # Example 3: Fill with a fixed constant
    # 'gear': lambda x: x.fillna(1),
}

Step 3: Apply Imputation Within Groups

Use groupby('trip') to split the data into trip groups, then apply a lambda function that uses fillna() with your strategy dictionary to each group:

# Apply the strategy to each group
df_imputed = df.groupby('trip', group_keys=False).apply(
    lambda group: group.fillna(impute_strategy)
)

The group_keys=False parameter keeps the output DataFrame's index clean (optional but helpful for readability).

Step 4: Verify the Result

Let's check the output for trip 1 to confirm the imputation works:

print(df_imputed[df_imputed['trip'] == 1])

You'll see:

  • brake values are forward-filled (rows 2-3 get 1.0 from row 1, rows 5-6 get 0.0 from row 4)
  • speed values are interpolated (rows 0-1 get values between NaN and 1.264, row 4 gets value between 0.000 and 1.264, etc.)

Key Notes

  • Flexibility: The impute_strategy dictionary lets you define unique logic for every column—you can use any Pandas fill method (ffill, bfill, interpolate with different methods like spline or pad, or even custom functions).
  • Performance: For very large datasets, this approach is efficient because it processes each group independently and leverages Pandas' optimized methods.
  • Unspecified Columns: Any columns not listed in impute_strategy will remain unchanged (their missing values stay as NaN, if you want to fill them with a default, you can add a catch-all or run a separate fill step).

内容的提问来源于stack exchange,提问作者retorquere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:31:16