如何修改Pandas代码实现符合特定条件的Hospital列行重命名
Fixing the Hospital Renaming Logic for Empty GeneralRepresentation Values
Hey Andrew, great question! The issue with your original code is that it doesn’t account for rows where GeneralRepresentation is empty—it treats those empty values as part of the unique count and factorization, leading to unnecessary renaming for hospitals that only have empty entries or a mix of empty and one non-empty value. Let’s adjust the code to meet all your requirements:
Step-by-Step Modified Solution
First, let’s break down the fixes we need:
- Exclude empty
GeneralRepresentationvalues when calculating unique counts and generating suffixes - Explicitly preserve the original hospital name for any row where
GeneralRepresentationis empty
Here’s the updated code:
import pandas as pd import numpy as np # Sample test data to verify behavior data = { 'Hospital': ['A', 'A', 'B', 'B', 'C', 'C', 'D', 'D'], 'GeneralRepresentation': ['X', 'Y', 'Z', 'Z', np.nan, np.nan, 'W', np.nan] } df = pd.DataFrame(data) # 1. Create a mask for empty GeneralRepresentation values mask_empty = df['GeneralRepresentation'].isna() # 2. Group by Hospital, calculate metrics excluding empty values g = df.groupby('Hospital')['GeneralRepresentation'] # Count unique NON-EMPTY values per Hospital s2 = g.transform(lambda x: x.dropna().nunique()) # Factorize NON-EMPTY values (assign unique numbers), set empty positions to NaN s1 = g.transform(lambda x: pd.Series(pd.factorize(x.dropna())[0] + 1, index=x.dropna().index).reindex(x.index)) # 3. Apply the renaming logic with all conditions df['Hospital'] = np.where( mask_empty, # First priority: empty values keep original name df['Hospital'], np.where( s2 == 1, # Second: non-empty but only one unique value, keep original df['Hospital'], df['Hospital'] + '_' + s1.astype(str) # Third: multiple unique values, add suffix ) ) print(df)
Output Explanation
Running this on the sample data will give you:
Hospital GeneralRepresentation 0 A_1 X 1 A_2 Y 2 B Z 3 B Z 4 C NaN 5 C NaN 6 D W 7 D NaN
- Hospital A: Has two different non-empty values → renamed with suffixes
- Hospital B: Same non-empty value → no rename
- Hospital C: All empty values → no rename
- Hospital D: One non-empty value + one empty → no rename (since unique non-empty count is 1)
Key Fixes from Original Code
- We added
mask_emptyto explicitly handle rows with emptyGeneralRepresentation - When calculating
s2(unique count), we usex.dropna().nunique()to ignore empty values - For
s1(factorized suffixes), we only apply factorization to non-empty entries and reindex to keep empty positions as NaN (which won’t be used thanks to the mask) - The nested
np.whereensures we prioritize empty values first, then handle the unique count condition for non-empty rows
内容的提问来源于stack exchange,提问作者Andrew
相关产品推荐
相关产品推荐

