You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame的每一列创建值频率列?

为Pandas DataFrame的每一列添加对应频率列的正确实现方法

问题背景

原始DataFrame:

colorsanimals
yellowcat
yellowcat
redcat
redcat
bluecat

需要为每一列新增对应频率列,展示各值的出现频率,目标效果:

colorscolors_frequencyanimalsanimals_frequency
yellow40%cat100%
yellow40%cat100%
red40%cat100%
red40%cat100%
blue20%cat100%

尝试以下代码后,df.info()显示animals_frequency列无有效数据:

frequency = list()
for column in df.columns:
     series = (df[column].value_counts(normalize=True, dropna=True)*100)
     overview.append(series)

#overview list
o_colors = overview[0] 
o_animals = overview[1]

df['animals_frequency'] = o_animals

正确实现方法

代码失效原因是直接赋值value_counts返回的Series时,索引不匹配导致无法映射到原DataFrame的行。可以通过map()方法完成值与频率的对应,再格式化输出,完整代码如下:

import pandas as pd

# 构造原始数据
data = {
    'colors': ['yellow', 'yellow', 'red', 'red', 'blue'],
    'animals': ['cat', 'cat', 'cat', 'cat', 'cat']
}
df = pd.DataFrame(data)

# 遍历每列生成频率列
for col in df.columns:
    # 计算各值的百分比频率
    freq_series = df[col].value_counts(normalize=True, dropna=True) * 100
    # 映射频率并格式化为百分比字符串
    df[f"{col}_frequency"] = df[col].map(freq_series).apply(lambda x: f"{x:.0f}%")

print(df)

核心逻辑说明

  • value_counts(normalize=True)返回各值的占比(0-1区间),乘以100转为百分比数值
  • map(freq_series)根据原列的值匹配对应的频率,确保每一行都能获取到对应值的频率
  • apply(lambda x: f"{x:.0f}%")将数值格式化为整数百分比字符串,符合目标格式要求

内容的提问来源于stack exchange,提问作者Luciana Coelho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 19:20:47