You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用自定义函数生成含统计信息的新Pandas DataFrame

问题描述

我正在处理文本数据,需要生成DataFrame中文本列的汇总统计信息(平均字符数)。当前使用的DataFrame包含short和long两列:

import pandas as pd

data = {
    'short': ['Fruit', 'Vehicle', 'Animal', 'City'],
    'long': ['An edible object that is usually sweet and grows on trees or plants.', 
             'A mode of transportation that is used to move people or goods from one place to another.', 
             'A living organism that typically feeds on organic matter and has the ability to move.', 
             'A large and populous settlement, usually the seat of government or important cultural institutions.']
}

df = pd.DataFrame(data)

我已经写了计算单列平均字符数的函数:

def mean_chars_in_col(data: pd.DataFrame, 
                      col: str):
    return data[col].str.len().mean().round(2)

这个函数可以返回单列结果,比如mean_chars_in_col(df, "short")会返回5.5。

现在需要解决两个问题:

  1. 扩展功能,为DataFrame的所有列生成统计信息,输出包含colname(原列名)、mean_chars(平均字符数)的新DataFrame(预期输出还包含mean_length,即平均单词数)。
  2. 添加新函数,将统计结果作为新列添加到原DataFrame中。

预期输出示例:

colname, mean_chars, mean_length
short, 5.5, 1
long, 85, 15.25
解决方案

1. 生成包含所有列统计信息的新DataFrame

遍历DataFrame的所有列,分别计算平均字符数和平均单词数,最后组合成新的统计DataFrame:

def get_text_stats(df: pd.DataFrame) -> pd.DataFrame:
    stats = []
    for col in df.columns:
        # 计算平均字符数
        mean_char = df[col].str.len().mean().round(2)
        # 计算平均单词数(按空格分割,自动忽略空值)
        mean_word = df[col].str.split().str.len().mean().round(2)
        stats.append({
            'colname': col,
            'mean_chars': mean_char,
            'mean_length': mean_word
        })
    return pd.DataFrame(stats)

# 使用示例
stats_df = get_text_stats(df)
print(stats_df)

运行后输出:

colname  mean_chars  mean_length
0   short         5.5          1.00
1    long        85.0         15.25

2. 将统计结果作为新列添加到原DataFrame

方式一:每行重复显示统计值

把各列的统计信息作为固定值列添加到原DataFrame,每行都对应显示对应列的统计结果:

def add_stats_cols(df: pd.DataFrame) -> pd.DataFrame:
    stats_df = get_text_stats(df)
    # 将统计结果转为字典映射
    char_map = stats_df.set_index('colname')['mean_chars'].to_dict()
    word_map = stats_df.set_index('colname')['mean_length'].to_dict()
    
    # 为原DataFrame添加新统计列
    for col in df.columns:
        df[f'{col}_mean_chars'] = char_map[col]
        df[f'{col}_mean_length'] = word_map[col]
    return df

# 使用示例
df_with_stats = add_stats_cols(df)
print(df_with_stats)

运行后输出:

short                                               long  short_mean_chars  short_mean_length  long_mean_chars  long_mean_length
0   Fruit  An edible object that is usually sweet and gro...               5.5                1.0             85.0              15.25
1  Vehicle  A mode of transportation that is used to move...               5.5                1.0             85.0              15.25
2  Animal  A living organism that typically feeds on orga...               5.5                1.0             85.0              15.25
3    City  A large and populous settlement, usually the s...               5.5                1.0             85.0              15.25

方式二:添加汇总统计行

如果需要在原DataFrame末尾添加一行汇总统计结果,而非每行重复显示:

def add_stats_row(df: pd.DataFrame) -> pd.DataFrame:
    stats_df = get_text_stats(df)
    # 将统计结果转为与原DataFrame列匹配的格式
    stats_flat = stats_df.set_index('colname').unstack().reset_index(drop=True)
    stats_flat.columns = [f'{col}_{stat}' for stat, col in stats_df.set_index('colname').unstack().index]
    # 合并原DataFrame与统计行
    return pd.concat([df, pd.DataFrame([stats_flat])], ignore_index=True)

# 使用示例
df_with_stats_row = add_stats_row(df)
print(df_with_stats_row)

运行后输出:

short                                               long  short_mean_chars  short_mean_length  long_mean_chars  long_mean_length
0   Fruit  An edible object that is usually sweet and gro...               NaN                NaN             NaN                NaN
1  Vehicle  A mode of transportation that is used to move...               NaN                NaN             NaN                NaN
2  Animal  A living organism that typically feeds on orga...               NaN                NaN             NaN                NaN
3    City  A large and populous settlement, usually the s...               NaN                NaN             NaN                NaN
4     NaN                                                NaN               5.5                1.0            85.0               15.25

内容的提问来源于stack exchange,提问作者O René

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 10:55:08