You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas基于双人关联数据表生成三人组合的总和统计结果

实现方案

核心思路

  1. 先将原始两两人员的分值数据转换为可快速查询的映射结构,消除两个人的顺序影响
  2. 生成所有不重复的3人组合
  3. 对每个3人组合,取出3组两两对应分值相加得到总分

完整可运行代码

import pandas as pd
from itertools import combinations

# 1. 构造原始数据(你可以替换为自己的真实数据读取逻辑,比如pd.read_csv)
data = [
    ["Adam", "Adam", 0.00],
    ["Adam", "Jack", 5.78],
    ["Adam", "John", 9.43],
    ["Adam", "Wills", 33.78],
    ["Adam", "Robert", 13.43],
    ["Adam", "Mary", 23.95],
    ["Adam", "Jennifer", 7.48],
    ["Adam", "Patricia", 5.15],
    ["Jack", "Jack", 0.00],
    ["Jack", "John", 1.43],
    ["Jack", "Wills", 2.78],
    ["Jack", "Robert", 0.43],
    ["Jack", "Mary", 9.95],
    ["Jack", "Jennifer", 11.48],
    ["Jack", "Patricia", 15.15],
]
raw_df = pd.DataFrame(data, columns=["Person1", "Person2", "Total"])

# 2. 构建两两分值查询字典,排除自身配对的0值
score_map = {}
for _, row in raw_df.iterrows():
    p1, p2, total = row["Person1"], row["Person2"], row["Total"]
    if p1 == p2:
        continue
    # 对两人名称排序,保证(Adam,Jack)和(Jack,Adam)对应同一个key
    pair_key = tuple(sorted([p1, p2]))
    score_map[pair_key] = total

# 3. 提取所有不重复的人员名单
all_persons = pd.concat([raw_df["Person1"], raw_df["Person2"]]).unique()

# 4. 生成所有3人组合并计算总分
result_list = []
for p1, p2, p3 in combinations(all_persons, 3):
    # 分别获取三个两两组合的分值
    sum_total = score_map[tuple(sorted([p1, p2]))] + \
                score_map[tuple(sorted([p1, p3]))] + \
                score_map[tuple(sorted([p2, p3]))]
    # 保留两位小数和示例格式对齐
    result_list.append([p1, p2, p3, round(sum_total, 2)])

# 5. 转换为最终DataFrame
result_df = pd.DataFrame(result_list, columns=["Person1", "Person2", "Person3", "Total"])

效果验证

你示例的Adam、Jack、Mary组合,计算结果为 5.78 + 23.95 + 9.95 = 39.68,和预期完全一致。

大数据量优化方案

如果原始数据量超过10万行,可替换字典查询为Pandas向量化查询,性能提升更明显:

import numpy as np

# 构建排序后的索引表
sorted_df = raw_df.copy()
sorted_df[["Person1", "Person2"]] = pd.DataFrame(
    np.sort(sorted_df[["Person1", "Person2"]], axis=1),
    index=sorted_df.index
)
sorted_df = sorted_df.query("Person1 != Person2")\
                     .drop_duplicates(subset=["Person1", "Person2"])\
                     .set_index(["Person1", "Person2"])

# 后续查询直接用索引取值即可,比如:
# score = sorted_df.loc[("Adam", "Jack"), "Total"]

内容的提问来源于stack exchange,提问作者Light

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 23:39:03