Python:对DataFrame指定可选列按行求和并分组
Solution for Summing Optional Columns and Grouping
Hey, I’ve got you covered! Handling optional columns like E doesn’t have to be tricky—here’s a clean solution that works no matter if E is present in your DataFrame or not:
Step 1: Calculate Row-wise Sum of A, C, and E (if it exists)
The key here is to only include columns that actually exist in your DataFrame when calculating the sum. You can do this either with a quick list comprehension or using pandas' filter() method:
import pandas as pd import numpy as np # Recreate your sample data (both scenarios) # Case 1: DataFrame without E column df_no_E = pd.DataFrame(np.random.randint(0,10,size=(10, 4)), columns=list('ABCD')) x = np.array([[1,2]]) df_no_E['G'] = np.repeat(x,5) # Case 2: DataFrame with E column df_with_E = pd.DataFrame(np.random.randint(0,10,size=(10, 5)), columns=list('ABCDE')) df_with_E['G'] = np.repeat(x,5) # Define your target columns target_cols = ['A', 'C', 'E'] # Option 1: Use list comprehension to get existing columns existing_cols = [col for col in target_cols if col in df_no_E.columns] df_no_E['A_C_E_sum'] = df_no_E[existing_cols].sum(axis=1) # Option 2: Use filter() (shorter syntax, same result) df_with_E['A_C_E_sum'] = df_with_E.filter(items=target_cols).sum(axis=1)
Step 2: Perform Grouping Operation
Once you have the row-wise sum, you can group by column G and apply any aggregation you need (like sum, mean, count, etc.):
# Group by G and calculate total sum for each group (no E case) grouped_no_E = df_no_E.groupby('G')['A_C_E_sum'].sum() print("Grouped result without E:\n", grouped_no_E) # Group by G and calculate average for each group (with E case) grouped_with_E = df_with_E.groupby('G')['A_C_E_sum'].mean() print("\nGrouped result with E:\n", grouped_with_E)
Why This Works
- The list comprehension or
filter()method automatically ignores columns that aren’t present in the DataFrame, so you don’t have to write messy conditional checks for E’s existence. - This approach is flexible—you can easily add or remove target columns later by just updating the
target_colslist.
内容的提问来源于stack exchange,提问作者Justin
相关产品推荐
相关产品推荐

