You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中分组计算均值与标准差并绘制误差棒柱状图?

解决Pandas分组计算均值与标准差的问题

Hey there! Let's break this down for you—you don't need clunky loops for this, especially with a large DataFrame. Pandas has built-in tools that are way more efficient, but I'll cover both the recommended approach and a loop-based method just in case you need it.


首先:推荐用groupby(高效,适合大型数据集)

手动循环在处理大型DataFrame时会拖慢速度,Pandas的groupby是优化过的矢量化操作,不仅代码简洁,运行效率也高很多。下面分几种常见场景说明:

场景1:筛选包含特定字符串的行,计算CRPS和Age的均值/标准差

如果你的需求是先找出所有包含特定字符串的行(比如某列"Category"里有"your_target_str"),再计算这些行中CRPS和Age的统计量:

import pandas as pd

# 筛选包含目标字符串的行(na=False避免空值干扰)
filtered_rows = table[table["Category"].str.contains("your_target_str", na=False)]

# 同时计算CRPS和Age的均值与标准差
summary_stats = filtered_rows[["CRPS", "Age"]].agg(["mean", "std"])
print(summary_stats)

场景2:按「是否包含特定字符串」分组,计算每组的统计量

如果你想对比包含特定字符串的行和不包含的行的CRPS、Age均值/标准差:

# 新增一列标记是否包含目标字符串
table["has_target"] = table["Category"].str.contains("your_target_str", na=False)

# 按这个标记分组,计算CRPS和Age的均值、标准差
grouped_stats = table.groupby("has_target")[["CRPS", "Age"]].agg(["mean", "std"])
print(grouped_stats)

场景3:按CRPS分组(仅包含特定字符串的组),计算Age的统计量

如果你的需求是针对CRPS列中包含特定字符串的分组,计算对应Age的均值和标准差:

# 先把CRPS转成字符串(避免数值类型报错),标记是否包含目标字符串
table["crps_matches"] = table["CRPS"].astype(str).str.contains("your_target_str", na=False)

# 分组计算Age的均值和标准差
age_group_stats = table.groupby("crps_matches")["Age"].agg(["mean", "std"])
print(age_group_stats)

如果一定要用循环实现(不推荐大型数据集)

如果你确实需要用循环来完成(比如有特殊自定义逻辑),可以这样写:

# 先获取所有符合条件的分组值(比如从"Category"列中筛选包含目标字符串的唯一值)
target_groups = [g for g in table["Category"].unique() if "your_target_str" in str(g)]

# 初始化字典存储结果
result_dict = {}

# 循环每个目标分组
for group in target_groups:
    # 筛选当前分组的行
    current_data = table[table["Category"] == group]
    # 计算统计量
    stats = {
        "CRPS_mean": current_data["CRPS"].mean(),
        "CRPS_std": current_data["CRPS"].std(),
        "Age_mean": current_data["Age"].mean(),
        "Age_std": current_data["Age"].std()
    }
    result_dict[group] = stats

# 转成DataFrame方便后续绘图
result_df = pd.DataFrame.from_dict(result_dict, orient="index")
print(result_df)

绘制带标准差误差棒的柱状图

拿到统计结果后,用seaborn或matplotlib就能轻松画出误差棒柱状图:

import seaborn as sns
import matplotlib.pyplot as plt

# 示例1:用原始数据直接绘图(seaborn自动计算统计量)
sns.barplot(
    x="has_target", 
    y="Age", 
    data=table, 
    ci="sd"  # ci="sd"表示用标准差作为误差棒
)
plt.title("Age Mean with Std Error Bars")
plt.show()

# 示例2:用我们计算好的分组统计结果绘图
sns.barplot(
    x=result_df.index, 
    y="Age_mean", 
    yerr=result_df["Age_std"],  # 指定标准差作为误差棒
    data=result_df
)
plt.xticks(rotation=45)  # 旋转x轴标签避免重叠
plt.title("Age Mean by Target Groups (with Std Error Bars)")
plt.show()

内容的提问来源于stack exchange,提问作者florence-y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:35:29