You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将pandas DataFrame多值列转换为两两关联对用于Networkx网络分析

解决方法

核心思路是使用itertools.combinations生成指定长度的无序不重复组合,刚好满足同项目下人员两两配对、无反向重复的需求。

完整实现代码

写法1:易读的循环实现

import pandas as pd
from itertools import combinations

# 构造原始示例数据,你可以替换为你自己的DataFrame读取逻辑
df = pd.DataFrame({
    'item': ['a', 'b', 'c'],
    'names': ['moriz, jon, cate', 'jon, lenard', 'cate, martin, leo, jil']
})

# 拆分姓名字符串为列表
df['name_list'] = df['names'].str.split(', ')

result = []
for _, row in df.iterrows():
    # 遍历每个项目,生成所有2人组合
    for person1, person2 in combinations(row['name_list'], 2):
        result.append({
            'item': row['item'],
            'person 1': person1,
            'person 2': person2
        })

# 转为DataFrame格式
result_df = pd.DataFrame(result)

写法2:Pandas风格的无循环实现

如果你偏好更简洁的pandas语法,可以用以下写法得到相同结果:

import pandas as pd
from itertools import combinations

df = pd.DataFrame({
    'item': ['a', 'b', 'c'],
    'names': ['moriz, jon, cate', 'jon, lenard', 'cate, martin, leo, jil']
})
df['name_list'] = df['names'].str.split(', ')

result_df = df.explode('name_list')\
              .groupby('item')['name_list']\
              .apply(lambda x: pd.DataFrame(combinations(x, 2), columns=['person 1', 'person 2']))\
              .reset_index(level=1, drop=True)\
              .reset_index()

注意事项

  1. 如果原始姓名数据存在多余空格、空值的情况,可以在拆分姓名列表时增加清洗逻辑,避免生成无效配对:
df['name_list'] = df['names'].str.split(',')\
                            .apply(lambda x: [i.strip() for i in x if i.strip()])
  1. 上述代码生成的配对顺序和姓名在原单元格中的顺序保持一致,如果需要匹配自定义的配对顺序,可以对拆分后的name_list做排序、反转等调整,不会影响配对的有效性。

内容的提问来源于stack exchange,提问作者Tsiss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 10:54:03