基于Python 3 pandas实现薪酬按频次标准化为年度薪酬的问题
Hey there, the issue with your code is almost certainly due to chained assignment—a common pitfall in pandas where you might be modifying a copy of your DataFrame instead of the original, leading to no visible changes (or a SettingWithCopyWarning you might have missed). Let's fix this with a few reliable approaches:
Approach 1: Use loc for Safe Assignment (Most Recommended)
Pandas' loc method ensures you're modifying the original DataFrame directly, avoiding ambiguous chain indexing:
# Initialize the new column with Yearly values first df['NormalizedAnnualCompensation'] = df['CompTotal'] # Update Monthly values using loc df.loc[df['CompFreq'] == "Monthly", 'NormalizedAnnualCompensation'] = df['CompTotal'] * 12 # Update Weekly values using loc df.loc[df['CompFreq'] == "Weekly", 'NormalizedAnnualCompensation'] = df['CompTotal'] * 52
Approach 2: Use np.select for Clear Conditional Logic
If you prefer a more concise, single-step approach, numpy.select lets you define all conditions and corresponding values at once:
import numpy as np # Define your conditions and matching values conditions = [ df['CompFreq'] == "Monthly", df['CompFreq'] == "Weekly", df['CompFreq'] == "Yearly" ] values = [ df['CompTotal'] * 12, df['CompTotal'] * 52, df['CompTotal'] ] # Assign the normalized values in one line df['NormalizedAnnualCompensation'] = np.select(conditions, values, default=df['CompTotal'])
Approach 3: Use a Mapping Dictionary for Cleanest Code
This method leverages a dictionary to map frequency types to their multipliers, making the logic super readable and efficient:
# Create a multiplier dictionary for each frequency type freq_multiplier = { 'Yearly': 1, 'Monthly': 12, 'Weekly': 52 } # Multiply CompTotal by the mapped multiplier df['NormalizedAnnualCompensation'] = df['CompTotal'] * df['CompFreq'].map(freq_multiplier)
Why Your Original Code Failed
Your original code uses chained indexing like df['NormalizedAnnualCompensation'][df['CompFreq'] == "Monthly"]. Pandas can't guarantee this operates on the original DataFrame—it might be modifying a temporary copy instead, which is why you don't see changes. Using loc, np.select, or the mapping approach avoids this ambiguity.
内容的提问来源于stack exchange,提问作者Jim Cooper

