You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量对Pandas DataFrame应用binom_test并解决长度错误

问题解决:批量对Pandas DataFrame应用二项式检验并修复ValueError

错误原因

你调用binom_test(get_rescued(breed), get_total(breed), 0.08)时触发错误,是因为breed是整个DataFrame的列对象,传给get_rescued后返回的是所有品种的rescued值组成的Series。而scipy.stats.binom_test要求参数x是单个成功次数,或是与n长度严格匹配的数组,这里传入多值Series不符合要求,触发了长度错误。

批量处理方案

不需要逐个调用函数,直接利用Pandas的apply或分组聚合功能,就能一次性完成所有品种的二项式检验。

场景1:每行对应一个独立品种

如果你的DataFrame中每行对应一个唯一品种,直接对每行应用检验:

from scipy.stats import binom_test

# 新增p_value列,存储每个品种的二项式检验结果
df['p_value'] = df.apply(lambda row: binom_test(row['rescued'], row['total'], 0.08), axis=1)

# 查看结果(保留关键列)
print(df[['breed', 'rescued', 'total', 'p_value']])

场景2:存在重复品种记录

如果同一个品种有多条记录,先按品种分组求和,再计算检验结果:

from scipy.stats import binom_test

# 按品种分组,汇总救助数和总数
breed_summary = df.groupby('breed').agg(
    rescued=('rescued', 'sum'),
    total=('total', 'sum')
).reset_index()

# 批量计算每个品种的p值
breed_summary['p_value'] = breed_summary.apply(lambda row: binom_test(row['rescued'], row['total'], 0.08), axis=1)

print(breed_summary)

额外优化

你之前写的get_attribute等函数完全可以省略,直接通过DataFrame的行或分组操作获取对应数值,代码更简洁高效。

内容的提问来源于stack exchange,提问作者Diastat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 00:45:49