如何用Pandas计算排序任务分组后的标签精准度统计
解决方案:计算段落关联预测的精准度
咱们已经有了按paragraphA升序、prediction降序排好序的DataFrame sorted_grouped,接下来一步步完成统计和精准度计算:
步骤1:给每个段落组标记真实正样本总数
首先要给每个paragraphA对应的所有行,添加上该组内label=1的总数量,这样后续筛选前x条时更方便:
# 用transform把每组的真实正样本数映射到每一行 sorted_grouped['true_pos_count'] = sorted_grouped.groupby('paragraphA')['label'].transform('sum')
比如示例里Paragraph1的所有行,true_pos_count都会被设为3,和它的真实正样本总数一致。
步骤2:统计每组前x条里的真实正样本数
接下来对每个paragraphA组,只取前true_pos_count条记录,然后统计其中label=1的数量:
def count_retrieved_positives(group): # 拿到当前组的真实正样本数x x = group['true_pos_count'].iloc[0] # 取前x条,统计其中label=1的数量 return group.head(x)['label'].sum() # 分组处理得到每个组的检索正样本数 retrieved_stats = sorted_grouped.groupby('paragraphA').apply(count_retrieved_positives).reset_index(name='retrieved_true_pos') # 把真实正样本数也合并进来,方便后续计算 true_stats = sorted_grouped.groupby('paragraphA')['label'].sum().reset_index(name='true_pos_count') final_stats = retrieved_stats.merge(true_stats, on='paragraphA')
现在final_stats里就有每个paragraphA的真实正样本总数,以及前x条结果里的真实正样本数——比如示例里Paragraph1对应的true_pos_count=3,retrieved_true_pos=2。
步骤3:计算整体精准度
最后用所有组的检索真实正样本数总和,除以真实正样本数总和,得到整体精准度:
precision = final_stats['retrieved_true_pos'].sum() / final_stats['true_pos_count'].sum() print(f"段落关联预测的精准度:{precision:.4f}")
完整整合代码
如果想简化步骤,也可以把逻辑整合到一个分组处理函数里:
# 步骤1:添加真实正样本数字段 sorted_grouped['true_pos_count'] = sorted_grouped.groupby('paragraphA')['label'].transform('sum') # 步骤2:分组计算所需统计值 def process_single_group(group): x = group['true_pos_count'].iloc[0] return pd.Series({ 'true_pos_count': x, 'retrieved_true_pos': group.head(x)['label'].sum() }) final_stats = sorted_grouped.groupby('paragraphA').apply(process_single_group).reset_index() # 步骤3:计算精准度 precision = final_stats['retrieved_true_pos'].sum() / final_stats['true_pos_count'].sum() print(f"整体精准度:{precision:.4f}")
这样就能完全满足你的统计需求,逻辑清晰也方便调试。
内容的提问来源于stack exchange,提问作者Mia
相关产品推荐
相关产品推荐

