You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas和Numpy按列通过均值、标准差转换为Z值

用 Pandas 和 NumPy 实现列数据的 Z 值转换

给定矩阵 matrix = [[1,2,3], [4,5,6], [7,8,9]],先构建含 a、b、c 列的数据表,再对每列计算 Z 值(公式:(数值 - 列均值)/列标准差)。以下是具体实现方法:

方法一:直接用 Pandas 计算

Pandas 可通过简单算术运算直接实现,同时支持控制标准差的计算方式(总体/样本):

import pandas as pd

# 构建数据表
matrix = [[1,2,3], [4,5,6], [7,8,9]]
df = pd.DataFrame(matrix, columns=['a', 'b', 'c'])

# 计算 Z 值(使用总体标准差,对应示例中的2.45)
z_scores_df = (df - df.mean()) / df.std(ddof=0)

print(z_scores_df)

输出结果:

a         b         c
0 -1.224745 -1.224745 -1.224745
1  0.000000  0.000000  0.000000
2  1.224745  1.224745  1.224745

注:ddof=0 表示计算总体标准差(对应示例中 a 列的标准差≈2.45,即 sqrt(6));若需使用样本标准差(除以 n-1),则去掉 ddof=0,此时 a 列标准差为3。

方法二:用 NumPy 手动计算

若需要更底层的控制,用 NumPy 实现步骤如下:

import numpy as np
import pandas as pd

matrix = [[1,2,3], [4,5,6], [7,8,9]]
arr = np.array(matrix)

# 计算每列的均值和总体标准差
col_means = np.mean(arr, axis=0)
col_stds = np.std(arr, axis=0, ddof=0)

# 广播计算 Z 值
z_scores_arr = (arr - col_means) / col_stds

# 转成 Pandas DataFrame
z_scores_df = pd.DataFrame(z_scores_arr, columns=['a', 'b', 'c'])

print(z_scores_df)

输出结果和方法一完全一致。

验证示例结果

以 a 列 [1,4,7] 为例:

  • 均值:(1+4+7)/3 = 4
  • 总体标准差:sqrt(((1-4)² + (4-4)² + (7-4)²)/3) = sqrt(18/3) = sqrt(6) ≈2.45
  • Z 值:(1-4)/2.45≈-1.22,(4-4)/2.45=0,(7-4)/2.45≈1.22,与输出结果完全匹配。

内容的提问来源于stack exchange,提问作者jtoyhh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 04:30:58