You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas describe()结果重塑为单行格式?机器学习特征统计需求

将Pandas describe()结果转为单行格式

嘿,这个需求我做特征工程的时候也碰到过!默认的describe()输出是按统计指标分行的,确实不太方便直接作为特征输入模型,转成单行、列名带上「特征+统计量」的形式就舒服多了。

这里给你两种实用的实现方式,第一种简洁直接,第二种更灵活可控:

方法一:快速生成完整单行结果

直接利用Pandas的unstack()方法展开多层结构,再重塑成单行:

import pandas as pd
import numpy as np

# 你的示例数据
df = pd.DataFrame({'A' : [1,np.nan,3], 'B' : [20,30,40]})

# 获取describe结果并转成单行格式
desc_result = df.describe()
single_row_df = pd.DataFrame(
    desc_result.unstack().values.reshape(1, -1),
    columns=[f"{col}_{stat}" for col in desc_result.columns for stat in desc_result.index]
)

print(single_row_df)

运行后输出会是这样的单行DataFrame:

A_count  A_mean    A_std  A_min  A_25%  A_50%  A_75%  A_max  B_count  B_mean  B_std  B_min  B_25%  B_50%  B_75%  B_max
0      2.0     2.0  1.41421    1.0    1.5    2.0    2.5    3.0      3.0    30.0   10.0   20.0   25.0   30.0   35.0   40.0

方法二:自定义保留特定统计量

如果不需要describe()输出的所有指标(比如只需要count、mean、std),可以先筛选行再处理:

# 指定需要保留的统计量
target_stats = ['count', 'mean', 'std']
desc_result = df.describe().loc[target_stats]

# 生成单行结果
single_row_df = pd.DataFrame(
    desc_result.unstack().values.reshape(1, -1),
    columns=[f"{col}_{stat}" for col in desc_result.columns for stat in target_stats]
)

print(single_row_df)

这次输出就只包含你指定的统计量了:

A_count  A_mean    A_std  B_count  B_mean  B_std
0      2.0     2.0  1.41421      3.0    30.0   10.0

原理其实很简单:unstack()会把原DataFrame的列(A、B)和行索引(count、mean等)转成多层索引的Series,我们把这些值重塑成一行,再用「特征名+统计量」的组合作为新列名,完美契合特征工程里的单行特征需求~

内容的提问来源于stack exchange,提问作者mrgloom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:42:33