Pandas按分组取非place=1的首个post值生成tpost列空值问题求解
问题原因分析
你原来的两种写法错误原因一致:都先过滤了place>1的行再做分组聚合,所有不存在place>1记录的分组(比如示例中id为202010145的分组)会被直接排除在分组结果外,赋值时这些分组无法匹配到对应值,最终tpost为空。
正确实现方案
方案1:易读的自定义分组逻辑写法
适合数据量不大、优先代码可读性的场景:
import pandas as pd # 原始数据集构造 data = {'id':['202006106','202006106','202006106','202007153','202007155','202009094', '202009094','202009095','202010143','202010143','202010145'], 'subject':['B_Math','B_Math','B_Math','B_Math','B_Math','B_Math', 'B_Math','B_Math','B_Math','B_Math','B_Math'], 'class':['Class_4','Class_4','Class_4','Class_4','Class_4','Class_4', 'Class_4','Class_4','Class_4','Class_4','Class_4'], 'hf':[1,1,1,1,1,1,1,1,1,1,1], 'place':[4,6,7,4,6,2,8,4,2,7,1], 'post':[0.048,0.045,0.043,0.042,0.040,0.038,0.037,0.036,0.034,0.033,0.065]} df = pd.DataFrame(data) # 自定义分组内取数逻辑 def get_group_tpost(group): # 优先取place>1的第一个post值 mask_gt1 = group['place'] > 1 if mask_gt1.any(): return group.loc[mask_gt1, 'post'].iloc[0] # 无符合条件记录则取place=1的post值 return group.loc[group['place'] == 1, 'post'].iloc[0] # 分组应用逻辑生成tpost df['tpost'] = df.groupby(['id', 'subject', 'class', 'hf'], group_keys=False).apply(get_group_tpost)
方案2:高性能向量化写法
适合数据量较大、需要提升运行效率的场景,避免自定义apply的性能损耗:
df['tpost'] = df.assign(priority=df['place'].eq(1).astype(int)) \ .sort_values(by=['priority', 'place']) \ .groupby(['id', 'subject', 'class', 'hf'])['post'] \ .transform('first')
结果验证
运行后id=202010145对应行的tpost值为0.065,其余分组的tpost均为组内第一个place>1对应的post值,完全符合需求。
内容的提问来源于stack exchange,提问作者ManOnTheMoon
相关产品推荐
相关产品推荐

