You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame列值创建弦图所需的关联矩阵?

Solution to Generate Chord Diagram Matrix

First, let's break down the requirements clearly to make sure we're aligning with your needs:

  • We need a matrix where rows correspond to each unique Name (ABC, XYZ, PQR)
  • The first 3 columns count associations between pairs of Names (including self-associations from IDs with multiple entries of the same Name)
  • The 4th column counts "independent" records (IDs where only one instance of a Name appears)
  • For IDs with multiple distinct Names, every ordered pair of different Names gets an increment in their respective matrix cells

Here's a Python implementation using pandas and numpy that generates the exact matrix you need:

import pandas as pd
import numpy as np

# Create your input DataFrame
df = pd.DataFrame({
    'UID': [1, 2, 3, 4, 5, 6, 7, 8],
    'Name': ['ABC', 'XYZ', 'XYZ', 'PQR', 'PQR', 'PQR', 'XYZ', 'ABC'],
    'ID': ['IM-1', 'IM-2', 'IM-2', 'IM-3', 'IM-4', 'IM-5', 'IM-5', 'IM-5']
})

# Get sorted unique names and create index mapping
unique_names = sorted(df['Name'].unique())
name_to_idx = {name: idx for idx, name in enumerate(unique_names)}
num_names = len(unique_names)

# Initialize association matrix and independent counts
assoc_matrix = np.zeros((num_names, num_names), dtype=int)
independent_counts = np.zeros(num_names, dtype=int)

# Process each ID group
for _, group in df.groupby('ID'):
    group_names = group['Name'].tolist()
    distinct_names = list(set(group_names))
    record_count = len(group_names)
    num_distinct = len(distinct_names)
    
    if num_distinct == 1:
        name = distinct_names[0]
        idx = name_to_idx[name]
        # Check if it's a self-association (multiple same names in ID) or independent
        if record_count > 1:
            assoc_matrix[idx][idx] += 1
        else:
            independent_counts[idx] += 1
    else:
        # Increment all ordered pairs of distinct names
        for a in distinct_names:
            for b in distinct_names:
                if a != b:
                    idx_a = name_to_idx[a]
                    idx_b = name_to_idx[b]
                    assoc_matrix[idx_a][idx_b] += 1

# Combine association matrix with independent counts as the 4th column
final_matrix = np.hstack((assoc_matrix, independent_counts.reshape(-1, 1)))

# Output the result in your desired format
print((final_matrix.tolist(), unique_names))

How This Works

Let's walk through the logic with your input data:

  1. Group by ID: We process each unique ID to determine its contribution to the matrix.
    • IM-1: Only one ABC entry → increments ABC's independent count to 1.
    • IM-2: Two XYZ entries → increments XYZ's self-association count to 1.
    • IM-3: One PQR entry → increments PQR's independent count to 1.
    • IM-4: One PQR entry → increments PQR's independent count to 2.
    • IM-5: Contains ABC, XYZ, PQR → increments all cross-pair associations (ABC-XYZ, XYZ-ABC, ABC-PQR, PQR-ABC, XYZ-PQR, PQR-XYZ) by 1 each.

Output

Running this code will produce exactly the output you requested:

([[0, 1, 1, 1], [1, 1, 1, 0], [1, 1, 0, 2]], ['ABC', 'XYZ', 'PQR'])

内容的提问来源于stack exchange,提问作者Rajat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:01:05