如何在Pandas DataFrame中分组计算均值与标准差并绘制误差棒柱状图?
解决Pandas分组计算均值与标准差的问题
Hey there! Let's break this down for you—you don't need clunky loops for this, especially with a large DataFrame. Pandas has built-in tools that are way more efficient, but I'll cover both the recommended approach and a loop-based method just in case you need it.
首先:推荐用groupby(高效,适合大型数据集)
手动循环在处理大型DataFrame时会拖慢速度,Pandas的groupby是优化过的矢量化操作,不仅代码简洁,运行效率也高很多。下面分几种常见场景说明:
场景1:筛选包含特定字符串的行,计算CRPS和Age的均值/标准差
如果你的需求是先找出所有包含特定字符串的行(比如某列"Category"里有"your_target_str"),再计算这些行中CRPS和Age的统计量:
import pandas as pd # 筛选包含目标字符串的行(na=False避免空值干扰) filtered_rows = table[table["Category"].str.contains("your_target_str", na=False)] # 同时计算CRPS和Age的均值与标准差 summary_stats = filtered_rows[["CRPS", "Age"]].agg(["mean", "std"]) print(summary_stats)
场景2:按「是否包含特定字符串」分组,计算每组的统计量
如果你想对比包含特定字符串的行和不包含的行的CRPS、Age均值/标准差:
# 新增一列标记是否包含目标字符串 table["has_target"] = table["Category"].str.contains("your_target_str", na=False) # 按这个标记分组,计算CRPS和Age的均值、标准差 grouped_stats = table.groupby("has_target")[["CRPS", "Age"]].agg(["mean", "std"]) print(grouped_stats)
场景3:按CRPS分组(仅包含特定字符串的组),计算Age的统计量
如果你的需求是针对CRPS列中包含特定字符串的分组,计算对应Age的均值和标准差:
# 先把CRPS转成字符串(避免数值类型报错),标记是否包含目标字符串 table["crps_matches"] = table["CRPS"].astype(str).str.contains("your_target_str", na=False) # 分组计算Age的均值和标准差 age_group_stats = table.groupby("crps_matches")["Age"].agg(["mean", "std"]) print(age_group_stats)
如果一定要用循环实现(不推荐大型数据集)
如果你确实需要用循环来完成(比如有特殊自定义逻辑),可以这样写:
# 先获取所有符合条件的分组值(比如从"Category"列中筛选包含目标字符串的唯一值) target_groups = [g for g in table["Category"].unique() if "your_target_str" in str(g)] # 初始化字典存储结果 result_dict = {} # 循环每个目标分组 for group in target_groups: # 筛选当前分组的行 current_data = table[table["Category"] == group] # 计算统计量 stats = { "CRPS_mean": current_data["CRPS"].mean(), "CRPS_std": current_data["CRPS"].std(), "Age_mean": current_data["Age"].mean(), "Age_std": current_data["Age"].std() } result_dict[group] = stats # 转成DataFrame方便后续绘图 result_df = pd.DataFrame.from_dict(result_dict, orient="index") print(result_df)
绘制带标准差误差棒的柱状图
拿到统计结果后,用seaborn或matplotlib就能轻松画出误差棒柱状图:
import seaborn as sns import matplotlib.pyplot as plt # 示例1:用原始数据直接绘图(seaborn自动计算统计量) sns.barplot( x="has_target", y="Age", data=table, ci="sd" # ci="sd"表示用标准差作为误差棒 ) plt.title("Age Mean with Std Error Bars") plt.show() # 示例2:用我们计算好的分组统计结果绘图 sns.barplot( x=result_df.index, y="Age_mean", yerr=result_df["Age_std"], # 指定标准差作为误差棒 data=result_df ) plt.xticks(rotation=45) # 旋转x轴标签避免重叠 plt.title("Age Mean by Target Groups (with Std Error Bars)") plt.show()
内容的提问来源于stack exchange,提问作者florence-y
相关产品推荐
相关产品推荐

