求助:Pandas遍历现有列并基于条件创建新列遇问题
Hey there! Let's work through your problem of creating multiple conditional columns (Month1 to Month4) in your DataFrame, and fix that warning you're seeing too.
First, let's tackle that warning you got
If you were getting a SettingWithCopyWarning when trying basic column creation, it's almost always because you're working with a slice of another DataFrame (like df = original_df[['RO_col1', 'RO_col2']] instead of making a copy). The fix is simple: always create a copy of your DataFrame first before modifying it:
# Make a copy to avoid warnings df = df.copy()
Or if you're starting from a slice:
df = original_df[['RO_col1', 'RO_col2', 'other_cols']].copy()
Creating multiple conditional columns efficiently
Instead of writing code for each column one by one (which gets repetitive), we can handle this in a clean, scalable way. Let's break it down based on what your conditions are:
Scenario 1: Each Month column uses a similar condition (e.g., different thresholds for RO_ columns)
Suppose your logic is:
- Month1: Mark "Yes" if any RO_ column is > 10, else "No"
- Month2: Mark "Yes" if any RO_ column is > 20, else "No"
- Month3: Mark "Yes" if any RO_ column is < 10, else "No"
- Month4: Mark "Yes" if all RO_ columns are > 10, else "No"
We can define these conditions in a dictionary and loop through them to create columns in one go:
import pandas as pd import numpy as np # Example DataFrame with RO_ prefix columns data = { 'RO_Sales': [12, 8, 25, 5], 'RO_Profit': [9, 22, 14, 30], 'RO_Expense': [18, 11, 7, 21] } df = pd.DataFrame(data).copy() # Remember the copy! # Define your month conditions month_rules = { 'Month1': lambda x: x.filter(like='RO_').gt(10).any(axis=1), 'Month2': lambda x: x.filter(like='RO_').gt(20).any(axis=1), 'Month3': lambda x: x.filter(like='RO_').lt(10).any(axis=1), 'Month4': lambda x: x.filter(like='RO_').gt(10).all(axis=1) } # Loop to create each column for col_name, condition in month_rules.items(): df[col_name] = np.where(condition(df), 'Yes', 'No')
Scenario 2: Each Month column uses unique, specific conditions
If each Month column needs to check a specific RO_ column or a unique combination of conditions, you can use df.assign() to create all columns in a single chained operation (no loops needed, and no warnings):
df = df.assign( # Month1: Check if RO_Sales > 15 Month1=np.where(df['RO_Sales'] > 15, 'Pass', 'Fail'), # Month2: Check if RO_Profit is between 10 and 25 Month2=np.where(df['RO_Profit'].between(10, 25), 'Pass', 'Fail'), # Month3: Check if RO_Expense < 10 OR RO_Sales > 20 Month3=np.where((df['RO_Expense'] < 10) | (df['RO_Sales'] > 20), 'Pass', 'Fail'), # Month4: Check if all RO_ columns are positive Month4=np.where((df['RO_Sales'] > 0) & (df['RO_Profit'] > 0) & (df['RO_Expense'] > 0), 'Pass', 'Fail') )
⚠️ Note: Use & for "and" and | for "or" in pandas boolean conditions, and wrap each condition in parentheses!
Why this works better than single-column creation
- Using loops or
assign()keeps your code DRY (Don't Repeat Yourself) - Making a copy of the DataFrame eliminates that annoying
SettingWithCopyWarning - You can easily adjust conditions by modifying the dictionary or
assign()arguments without rewriting redundant code
内容的提问来源于stack exchange,提问作者CandleWax

