在Python Pandas中按两列分组后将剩余聚合列转为列表
To achieve your desired result—grouping by BN and PN while aggregating the other columns into lists (or keeping scalars for single entries)—here's a step-by-step approach using Pandas:
Step 1: Define the Custom Aggregation Function
We need a function that returns a scalar if there's only one value in the group, otherwise returns a list of values:
import pandas as pd def agg_list_or_scalar(series): # Return scalar if single value, else list of values return series.iloc[0] if len(series) == 1 else series.tolist()
Step 2: Create Your Sample DataFrame
Let's replicate your input data first:
data = { 'BN': [7363311, 7363311, 7363311, 7363311, 7363311, 7363311, 7363311, 7363311], 'PN': [1, 2, 3, 4, 4, 5, 7, 7], 'tempC': [28, 27, 27, 27, 27, 27, 27, 27], 'tempF': [82, 81, 81, 81, 81, 81, 81, 81], 'humidity': [73, 73, 73, 73, 73, 73, 73, 74] } df = pd.DataFrame(data)
Step 3: Group and Aggregate
Use groupby on ['BN', 'PN'] and apply our custom function to the target columns:
# Group by BN and PN, then aggregate each column grouped_df = df.groupby(['BN', 'PN'], as_index=False).agg({ 'tempC': agg_list_or_scalar, 'tempF': agg_list_or_scalar, 'humidity': agg_list_or_scalar })
At this point, grouped_df contains the correct values (scalars for single entries, lists for duplicates), but the BN column repeats the same value for every row. To match your desired output formatting, we can hide duplicate BN values.
Step 4: Format for Display (Optional)
To make the BN column appear only once per group (like your example), replace duplicate values with empty strings:
# Replace duplicate BN values with empty strings for cleaner display grouped_df['BN'] = grouped_df['BN'].mask(grouped_df['BN'].duplicated(), '') # Print the formatted result print(grouped_df.to_string(index=False))
Output
Running the above code will produce:
BN PN tempC tempF humidity 7363311 1 28 82 73 2 27 81 73 3 27 81 73 4 [27, 27] [81, 81] [73, 73] 5 27 81 73 7 [27, 27] [81, 81] [73, 74]
Key Notes
- The custom aggregation function ensures that single-value groups stay as scalars instead of being wrapped in a list (which is what happens if you just use
agg(list)). - The display formatting step is optional—if you need the underlying data to keep all
BNvalues, skip that part. - If you want to keep
BNandPNas the index (instead of columns), omitas_index=Falsein thegroupbycall, but you'll need to adjust the display formatting accordingly.
内容的提问来源于stack exchange,提问作者Tarun Khaneja

