基于Stratified Sampler生成DataFrame平衡随机列的技术问询
分层生成平衡随机A/B列解决方案
问题背景
给定如下DataFrame:
import pandas as pd df = pd.DataFrame({ "x": [0, 0, 1, 1, 0, 0, 1, 1], "y": [1, 2, 1, 2, 2, 2, 1, 1], })
需要编写函数生成包含"A"和"B"的随机列,满足要求:
- 针对指定分层列子集(如仅
x,或x与y组合),每个分层组内"A""B"出现次数尽可能平衡 - 组内样本数为偶数时,两者次数完全相等;为奇数时,次数差值不超过1
以x为分层列的示例结果如下:
import pandas as pd df = pd.DataFrame({ "x": [0, 0, 1, 1, 0, 0, 1, 1], "y": [1, 2, 1, 2, 2, 2, 1, 1], "outcome": ["A", "B", "A", "B", "A", "B", "A", "B"] })
解决方案
函数实现
import pandas as pd import numpy as np def generate_balanced_ab(df, group_cols): # 生成单组内平衡的A/B序列 def create_group_ab(n): count_a = n // 2 count_b = n - count_a ab_list = ["A"] * count_a + ["B"] * count_b np.random.shuffle(ab_list) return ab_list # 按指定列分组生成序列并拼接 df["outcome"] = df.groupby(group_cols, group_keys=False).apply( lambda g: pd.Series(create_group_ab(len(g)), index=g.index) ) return df
示例使用
- 以
x作为分层列:
df = pd.DataFrame({ "x": [0, 0, 1, 1, 0, 0, 1, 1], "y": [1, 2, 1, 2, 2, 2, 1, 1], }) result_df = generate_balanced_ab(df, group_cols=["x"]) print(result_df)
输出示例(随机打乱后,每个x组内A/B数量完全相等):
x y outcome 0 0 1 A 1 0 2 B 2 1 1 A 3 1 2 B 4 0 2 B 5 0 2 A 6 1 1 B 7 1 1 A
- 以
x和y的组合作为分层列:
result_df = generate_balanced_ab(df, group_cols=["x", "y"]) print(result_df)
输出示例(每个(x,y)组内A/B数量平衡):
x y outcome 0 0 1 A # 组内仅1个样本,A/B随机分配 1 0 2 A 2 1 1 A 3 1 2 A # 组内仅1个样本,A/B随机分配 4 0 2 B 5 0 2 B 6 1 1 B 7 1 1 A
函数说明
create_group_ab(n):根据组内样本数n生成基础A/B列表,偶数时A/B数量完全相等,奇数时差值为1,最后随机打乱保证随机性groupby(group_cols, group_keys=False):按指定列分组,group_keys=False避免额外添加分组键索引apply:为每个分组应用序列生成函数,最终拼接为整列
内容的提问来源于stack exchange,提问作者David Masip
相关产品推荐
相关产品推荐

