You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为DataFrame的ID列值计算关联权重并添加权重列

为超大DataFrame按ID频率添加样本权重的高效方法

我正在处理一个超大DataFrame,样本数据如下:

import pandas as pd
import numpy as np

df = pd.DataFrame({ 
    'ID': ['A', 'A', 'A', 'X', 'X', 'Y'], 
})

对应的DataFrame内容:

ID
0  A
1  A
2  A
3  X
4  X
5  Y

需要基于ID列各值的出现频率,用以下自定义函数计算权重,再高效为每行添加对应ID的sample_weight列:

def get_weights_inverse_num_of_samples(label_counts, power=1.):
    no_of_classes = len(label_counts)
    weights_for_samples = 1.0/np.power(np.array(label_counts), power)
    weights_for_samples = weights_for_samples / np.sum(weights_for_samples) * no_of_classes
    return weights_for_samples

# 计算ID的出现频率
freq = df['ID'].value_counts()
print(freq)

输出的频率统计:

ID
A    3
X    2
Y    1
Name: count, dtype: int64
# 计算各ID对应的权重
weights = get_weights_inverse_num_of_samples(freq)
print(weights)

输出的权重数组:

[0.54545455 0.81818182 1.63636364]

高效添加权重列到原DataFrame

针对超大DataFrame,要避免低效循环,我们可以先构建ID到权重的映射字典,再用map方法快速赋值:

# 构建ID与权重的映射关系
weight_map = dict(zip(freq.index, weights))
# 添加sample_weight列
df['sample_weight'] = df['ID'].map(weight_map)

最终得到的DataFrame:

ID  sample_weight
0  A        0.545455
1  A        0.545455
2  A        0.545455
3  X        0.818182
4  X        0.818182
5  Y        1.636364

内容的提问来源于stack exchange,提问作者armin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:35:23