You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将全成对距离矩阵转为无重复成对距离列表及大矩阵性能分析

对称距离矩阵转上三角无对角线距离列表方案

1. 中小矩阵的pandas实现

直接用stack()结合索引过滤即可剔除重复值和对角线0:

import pandas as pd

# 构造示例对称矩阵
data = [
    [0.000, 1.154, 1.235, 1.297, 0.960],
    [1.154, 0.000, 0.932, 0.929, 0.988],
    [1.235, 0.932, 0.000, 0.727, 1.244],
    [1.297, 0.929, 0.727, 0.000, 1.019],
    [0.960, 0.988, 1.244, 1.019, 0.000]
]
df = pd.DataFrame(data, index=[1,2,3,4,5], columns=[1,2,3,4,5])

# 堆叠矩阵并过滤上三角(不含对角线)
stacked_series = df.stack()
filtered = stacked_series[stacked_series.index.get_level_values(0) < stacked_series.index.get_level_values(1)]

# 转换为三列格式
result_df = filtered.reset_index()
result_df.columns = ["item1", "item2", "distance"]

print(result_df)

执行后输出符合要求的结果:

item1  item2  distance
0      1      2     1.154
1      1      3     1.235
2      1      4     1.297
3      1      5     0.960
4      2      3     0.932
5      2      4     0.929
6      2      5     0.988
7      3      4     0.727
8      3      5     1.244
9      4      5     1.019

2. 100,000×100,000规模矩阵的优化方案

10万级别的矩阵元素量达1e10,直接用pandas会内存爆炸,必须用numpy或分块处理:

方案A:numpy上三角索引提取

利用np.triu_indices_from()直接获取上三角(跳过对角线)的索引,内存效率远高于pandas:

import numpy as np

# 假设dist_matrix是100000×100000的numpy对称矩阵
dist_matrix = np.random.rand(100000, 100000)
dist_matrix = (dist_matrix + dist_matrix.T) / 2  # 构造对称矩阵
np.fill_diagonal(dist_matrix, 0)

# 获取上三角(k=1表示跳过对角线)的行、列索引
rows, cols = np.triu_indices_from(dist_matrix, k=1)
distances = dist_matrix[rows, cols]

# 组合为三列数组(索引从1开始则加1)
result_array = np.column_stack([rows + 1, cols + 1, distances])

注意:该方案需要约120GB内存(5e9个int64索引占80GB,5e9个float64距离占40GB),仅适用于大内存服务器。

方案B:分块处理(低内存场景)

将矩阵分块处理,边计算边写入磁盘,避免一次性加载所有数据:

import numpy as np
import csv

chunk_size = 1000  # 每次处理1000行,可根据内存调整
n = 100000

# 写入结果到CSV
with open('distance_list.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['item1', 'item2', 'distance'])
    
    for start_row in range(0, n, chunk_size):
        end_row = min(start_row + chunk_size, n)
        chunk = dist_matrix[start_row:end_row, :]
        
        for row_offset in range(chunk.shape[0]):
            global_row = start_row + row_offset
            # 只处理列索引大于当前行的部分
            target_cols = range(global_row + 1, n)
            row_distances = chunk[row_offset, target_cols]
            
            # 逐行写入
            for col, dist in zip(target_cols, row_distances):
                writer.writerow([global_row + 1, col + 1, dist])

该方案内存占用仅为chunk_size × n的矩阵内存(约800MB,float64),但处理时间较长,取决于磁盘IO速度。

3. 性能对比

  • pandas stack():仅适用于≤1000×1000的矩阵,1000级规模耗时几秒;10万级规模直接内存溢出,完全不可用。
  • numpy triu_indices:10万级规模在大内存服务器上耗时几十分钟,内存占用极高。
  • 分块处理:10万级规模耗时数小时,内存占用低,适合普通硬件。

内容的提问来源于stack exchange,提问作者Philipp O.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 03:07:10