You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于id分组计算多列Z-score并生成标签列

Pandas分组计算Z-score并生成标签列的实现方案

刚好做过类似的需求,这就给你一步步拆解实现方法,全程不用循环,高效又简洁:

1. 先初始化基础数据

首先把题目里给定的原始DataFrame创建出来:

import pandas as pd
import numpy as np
from scipy.stats import zscore

# 创建原始DataFrame
df = pd.DataFrame(
    [["A",1,98,56,61], ["B",1,99,54,36], ["C",1,97,32,83],
     ["B",1,96,31,90], ["C",1,45,32,12], ["A",1,67,33,55],
     ["C",1,54,65,73], ["A",1,34,84,98], ["B",1,76,12,99]],
    columns=["id","date","c1","c2","c3"]
)

2. 分组计算Z-score并新增列

我们利用Pandas的groupby+transform组合来实现分组内的Z-score计算,transform会自动把计算结果匹配回原DataFrame的对应行,完美保留原始数据结构:

# 指定需要处理的列
target_cols = ["c1", "c2", "c3"]

# 批量生成Z-score列
for col in target_cols:
    df[f"{col}_zscore"] = df.groupby('id')[col].transform(zscore)

如果不想依赖scipy,也可以手动计算Z-score,把上面的zscore替换成lambda x: (x - x.mean())/x.std(),效果完全一致。

3. 生成标签列

用np.where做矢量化判断,一次性生成所有标签列,比循环高效太多:

# 批量生成标签列:Z-score < -1标记为-1,其余为1
for col in target_cols:
    df[f"{col}_tag"] = np.where(df[f"{col}_zscore"] < -1, -1, 1)

4. 验证结果

运行完上面的代码后,你的DataFrame就和题目里给出的df_out完全一致了。如果要精确对比,可以用Pandas的断言工具(注意浮点精度的微小差异不影响结果):

# 假设df_out是题目里的预期结果,执行以下代码验证
pd.testing.assert_frame_equal(df.round(6), df_out.round(6), check_dtype=False)

内容的提问来源于stack exchange,提问作者Chethan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:03:14