如何基于指定列组的存在性创建DataFrame新列
I get it, dealing with conditional column creation based on missing columns can get messy with loops if you don’t structure it right. Here’s a clean, maintainable approach that uses a configuration dictionary to define your groups and their aggregation rules, then iterates through them to check for column presence:
First, define your group configurations in a dict where each key is the name of the new column you want to create, and the value is a tuple of (list of columns in the group, aggregation function to apply row-wise):
import pandas as pd # Define group configurations: {new_column_name: (column_list, agg_function)} group_configs = { 'Group1': (['A', 'B', 'D'], 'min'), 'Group2': (['C', 'E'], 'max') } # Sample DataFrame (replace this with your actual df) df = pd.DataFrame({ 'A': [1, 2, 3], 'B': [4, 5, 6], 'C': [7, 8, 9], 'D': [10, 11, 12], 'E': [13, 14, 15] })
Then, loop through each group in the config, check if all columns in the group exist in your DataFrame, and create the new column only if they do:
for new_col, (cols, agg_func) in group_configs.items(): # Check if all columns in the group are present in the DataFrame if set(cols).issubset(df.columns): df[new_col] = df[cols].agg(agg_func, axis=1)
Let’s verify this works with your scenarios:
- All columns present: The code creates both
Group1(row-wise min of A,B,D) andGroup2(row-wise max of C,E) as expected. - E column missing:
set(['C','E']).issubset(df.columns)returns False, soGroup2is skipped—onlyGroup1is created. - A and D missing:
set(['A','B','D']).issubset(df.columns)returns False, soGroup1is skipped—onlyGroup2is created. - A and C missing: Both group checks fail, so no new columns are added.
This approach is scalable too—if you need to add more groups later, just add another entry to the group_configs dict without modifying the loop logic.
内容的提问来源于stack exchange,提问作者Stan

