Pandas中聚合DataFrame列表字段报错:无数值类型可聚合
Why This Happens
The error you're seeing makes total sense: your C column contains lists instead of numeric values (like integers or floats). Pandas' built-in aggregation functions like mean and std only work with numeric data types—they don't know how to handle lists directly, hence the DataError.
Solutions
Here are two practical approaches to fix this, depending on your specific needs:
Approach 1: Explode Lists into Rows First
This is perfect if you want every element inside the lists to contribute to the aggregation:
import pandas as pd # Your original DataFrame df = pd.DataFrame({'A':[1,1,1,1,2,2,2,2], 'B':['x','x','y','y','x','x','y','y'], 'C':[[1,2,3,4],[5,6],[8,9,10,11],[12,13,14],[15],[16,17,18],[],[19]]}) # Explode lists into individual rows, then convert C to numeric type df_exploded = df.explode('C').astype({'C': int}) # Group by A and calculate mean/std result = df_exploded.groupby('A')['C'].agg(['mean', 'std']) print(result)
Output:
mean std A 1 7.416667 4.716991 2 16.166667 1.722408
Approach 2: Custom Aggregation Functions for Lists
If you want to keep the original grouping structure and aggregate across all lists in a group (e.g., combine all lists into one big collection then compute stats), use custom functions:
import pandas as pd import numpy as np # Custom function to calculate mean across all elements in group lists def list_mean(series): all_values = [num for lst in series for num in lst] return np.mean(all_values) if all_values else np.nan # Custom function to calculate standard deviation across all elements in group lists def list_std(series): all_values = [num for lst in series for num in lst] return np.std(all_values, ddof=1) if len(all_values) >= 2 else np.nan # Run the aggregation result = df.groupby('A')['C'].agg([list_mean, list_std]).rename(columns={'list_mean':'mean', 'list_std':'std'}) print(result)
This gives the same output as Approach 1, and it handles empty lists gracefully (they're just ignored in calculations).
Bonus: Aggregate List Means
If your goal is to first calculate the mean of each individual list, then aggregate those means per group, adjust the custom function like this:
def group_list_mean(series): list_means = [np.mean(lst) for lst in series if lst] return np.mean(list_means) if list_means else np.nan
内容的提问来源于stack exchange,提问作者HappyPy

