Pandas DataFrame列的条件填充问题咨询
Solution for Conditional Column Filling in Pandas DataFrame
Let's fix your code to meet the three specific rules you outlined. First, let's recap your requirements clearly to make sure we're aligned:
- Rule 1 (Code starts with "A"): Keep all original values in
NumxandNumyfor every row in the group. - Rule 2 (Code starts with "B"): Only keep the values from the first row of the group (the non-empty
Coderow), set all subsequent rows in the group toNaN. - Rule 3 (Code starts with "C"): Keep the first non-null float value in
Numx/Numy, and carry it forward until a new non-null float value appears in the group.
What Was Wrong with the Original Code?
Your initial code had two key issues:
- It only handled the
A1group explicitly, ignoring all other "A" prefix codes. - It didn't implement any logic for "B" or "C" prefix groups, which led to unintended
NaNvalues across those rows.
Step-by-Step Correct Implementation
First, let's preprocess the data to clean up types and create proper group identifiers:
import pandas as pd import numpy as np # Test data definition d = ({ 'Code' :['A1','A1','','B1','B1','A2','A2','','B2','B2','','A3','A3','A3','','B1','','B4','B4','A2','A2','A1','A1','','B4','B4','C1','C1','','','D1','','B2'], 'Numx' : [30.2,30.5,30.6,35.6,40.2,45.5,46.1,48.1,48.5,42.2,'',30.5,30.6,35.6,40.2,45.5,'',48.1,48.5,42.2, 40.1,48.5,42.2,'',48.5,42.2,43.1,44.1,'','','','',45.1], 'Numy' : [1.9,2.3,2.5,2.2,2.5,3.1,3.4,3.6,3.7,5.4,'',2.3,2.5,2.2,2.5,3.1,'',3.6,3.7,5.4,6.5,8.5,2.2,'',8.5,2.2,2.3,2.5,'','','','',3.2] }) df = pd.DataFrame(data=d) # Step 1: Convert Numx/Numy to numeric (empty strings become NaN) df[['Numx', 'Numy']] = df[['Numx', 'Numy']].apply(pd.to_numeric, errors='coerce') # Step 2: Create group identifier by forward-filling empty Code values df['group_code'] = df['Code'].ffill() # Step 3: Mark the start of each new group (non-empty Code rows) df['is_group_start'] = df['Code'] != '' # Step 4: Apply rules based on group_code prefix for col in ['Numx', 'Numy']: # Rule 1: A-prefix groups - keep all values (no change needed, but we'll explicitly retain) a_mask = df['group_code'].str.startswith('A') df.loc[a_mask, col] = df.loc[a_mask, col] # Rule 2: B-prefix groups - only keep values at group start, others NaN b_mask = df['group_code'].str.startswith('B') df.loc[b_mask & ~df['is_group_start'], col] = np.nan # Rule 3: C-prefix groups - forward-fill until new non-null value appears c_mask = df['group_code'].str.startswith('C') # Within C groups, forward fill only when the current value is NaN df.loc[c_mask, col] = df.loc[c_mask, col].ffill() # Optional: Drop helper columns if needed # df = df.drop(['group_code', 'is_group_start'], axis=1) print(df)
How This Works
- Type Conversion: We first convert
NumxandNumyto numeric types, turning empty strings intoNaNwhich Pandas can handle consistently. - Group Identification: Using
ffill()on theCodecolumn gives us a consistent group label for every row, even whenCodeis empty. - Group Start Marker: The
is_group_startcolumn helps us identify the first row of each group (the non-emptyCoderow) for B-prefix groups. - Rule Application:
- For A-prefix groups: We leave values as-is since we need to retain all original data.
- For B-prefix groups: We null out all values except the first row of the group.
- For C-prefix groups: We use
ffill()within the C groups to carry forward the last valid value until a new non-null value is encountered.
This implementation strictly follows all three rules you defined and avoids the excessive NaN values from your original code.
内容的提问来源于stack exchange,提问作者user9639519
相关产品推荐
相关产品推荐

