You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按MarketId和SelectionId计算DataFrame中Prob的均值?

问题:按MarketId和SelectionId分组计算Prob列均值

我尝试用以下代码分组计算Prob列的均值,但未得到预期结果:

df.groupby(['MarketId', 'SelectionId', ], as_index=False)['Prob'].mean()

示例DataFrame

TimeMarketIdSelectionIdProb
006/01/2016 19:58:011.12211769563433.3
106/01/2016 19:58:011.12211769479992.34
206/01/2016 19:58:011.12211769588053.8
306/01/2016 19:59:011.12211769563433.2
406/01/2016 19:59:011.12211769479992.3
506/01/2016 19:59:011.12211769588053.8
606/01/2016 20:00:011.12211769563433.2
706/01/2016 20:00:011.12211769479992.34
806/01/2016 20:00:011.12211769588053.8
915/06/2016 18:59:431.122271208241.25
1015/06/2016 18:59:431.1222712081528519
1115/06/2016 18:59:431.122271208588056.6
1215/06/2016 19:01:431.122271208241.26
1315/06/2016 19:01:431.1222712081528518
1415/06/2016 19:01:431.122271208588056.8
1515/06/2016 19:02:431.122271208241.27
1615/06/2016 19:02:431.1222712081528519
1715/06/2016 19:02:431.122271208588056.6

期望输出DataFrame

MarketIdSelectionIdProb
01.12211769563433.233
11.12211769479992.326
21.12211769588053.8
31.122271208241.26
41.1222712081528518.667
51.122271208588056.667

解决方案

你的代码逻辑本身是正确的,未得到预期结果通常是两个原因导致:

1. 浮点精度问题

MarketId为浮点数时,可能存在微小的存储精度差异(比如1.12211769实际存储为1.1221176900000001),导致本该同组的记录被错误拆分。可以将MarketId转为字符串类型避免这个问题:

# 转换MarketId为字符串,消除浮点精度影响
df['MarketId'] = df['MarketId'].astype(str)
# 确保SelectionId为整数类型,避免类型不一致导致的分组错误
df['SelectionId'] = df['SelectionId'].astype(int)

2. 小数位数格式问题

期望输出的均值保留了3位小数,而默认的mean()计算会返回更多小数位,需要用round(3)调整格式:

最终可用代码:

# 预处理数据类型
df['MarketId'] = df['MarketId'].astype(str)
df['SelectionId'] = df['SelectionId'].astype(int)

# 分组计算均值并保留3位小数
result = df.groupby(['MarketId', 'SelectionId'], as_index=False)['Prob'].mean().round(3)

执行后即可得到与期望一致的输出结果。

内容的提问来源于stack exchange,提问作者Robsmith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 22:16:17