You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame列的条件填充问题咨询

Solution for Conditional Column Filling in Pandas DataFrame

Let's fix your code to meet the three specific rules you outlined. First, let's recap your requirements clearly to make sure we're aligned:

  • Rule 1 (Code starts with "A"): Keep all original values in Numx and Numy for every row in the group.
  • Rule 2 (Code starts with "B"): Only keep the values from the first row of the group (the non-empty Code row), set all subsequent rows in the group to NaN.
  • Rule 3 (Code starts with "C"): Keep the first non-null float value in Numx/Numy, and carry it forward until a new non-null float value appears in the group.

What Was Wrong with the Original Code?

Your initial code had two key issues:

  1. It only handled the A1 group explicitly, ignoring all other "A" prefix codes.
  2. It didn't implement any logic for "B" or "C" prefix groups, which led to unintended NaN values across those rows.

Step-by-Step Correct Implementation

First, let's preprocess the data to clean up types and create proper group identifiers:

import pandas as pd
import numpy as np

# Test data definition
d = ({ 
    'Code' :['A1','A1','','B1','B1','A2','A2','','B2','B2','','A3','A3','A3','','B1','','B4','B4','A2','A2','A1','A1','','B4','B4','C1','C1','','','D1','','B2'], 
    'Numx' : [30.2,30.5,30.6,35.6,40.2,45.5,46.1,48.1,48.5,42.2,'',30.5,30.6,35.6,40.2,45.5,'',48.1,48.5,42.2, 40.1,48.5,42.2,'',48.5,42.2,43.1,44.1,'','','','',45.1], 
    'Numy' : [1.9,2.3,2.5,2.2,2.5,3.1,3.4,3.6,3.7,5.4,'',2.3,2.5,2.2,2.5,3.1,'',3.6,3.7,5.4,6.5,8.5,2.2,'',8.5,2.2,2.3,2.5,'','','','',3.2] 
})
df = pd.DataFrame(data=d)

# Step 1: Convert Numx/Numy to numeric (empty strings become NaN)
df[['Numx', 'Numy']] = df[['Numx', 'Numy']].apply(pd.to_numeric, errors='coerce')

# Step 2: Create group identifier by forward-filling empty Code values
df['group_code'] = df['Code'].ffill()

# Step 3: Mark the start of each new group (non-empty Code rows)
df['is_group_start'] = df['Code'] != ''

# Step 4: Apply rules based on group_code prefix
for col in ['Numx', 'Numy']:
    # Rule 1: A-prefix groups - keep all values (no change needed, but we'll explicitly retain)
    a_mask = df['group_code'].str.startswith('A')
    df.loc[a_mask, col] = df.loc[a_mask, col]
    
    # Rule 2: B-prefix groups - only keep values at group start, others NaN
    b_mask = df['group_code'].str.startswith('B')
    df.loc[b_mask & ~df['is_group_start'], col] = np.nan
    
    # Rule 3: C-prefix groups - forward-fill until new non-null value appears
    c_mask = df['group_code'].str.startswith('C')
    # Within C groups, forward fill only when the current value is NaN
    df.loc[c_mask, col] = df.loc[c_mask, col].ffill()

# Optional: Drop helper columns if needed
# df = df.drop(['group_code', 'is_group_start'], axis=1)

print(df)

How This Works

  1. Type Conversion: We first convert Numx and Numy to numeric types, turning empty strings into NaN which Pandas can handle consistently.
  2. Group Identification: Using ffill() on the Code column gives us a consistent group label for every row, even when Code is empty.
  3. Group Start Marker: The is_group_start column helps us identify the first row of each group (the non-empty Code row) for B-prefix groups.
  4. Rule Application:
    • For A-prefix groups: We leave values as-is since we need to retain all original data.
    • For B-prefix groups: We null out all values except the first row of the group.
    • For C-prefix groups: We use ffill() within the C groups to carry forward the last valid value until a new non-null value is encountered.

This implementation strictly follows all three rules you defined and avoids the excessive NaN values from your original code.

内容的提问来源于stack exchange,提问作者user9639519

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:52:43