如何将全成对距离矩阵转为无重复成对距离列表及大矩阵性能分析
对称距离矩阵转上三角无对角线距离列表方案
1. 中小矩阵的pandas实现
直接用stack()结合索引过滤即可剔除重复值和对角线0:
import pandas as pd # 构造示例对称矩阵 data = [ [0.000, 1.154, 1.235, 1.297, 0.960], [1.154, 0.000, 0.932, 0.929, 0.988], [1.235, 0.932, 0.000, 0.727, 1.244], [1.297, 0.929, 0.727, 0.000, 1.019], [0.960, 0.988, 1.244, 1.019, 0.000] ] df = pd.DataFrame(data, index=[1,2,3,4,5], columns=[1,2,3,4,5]) # 堆叠矩阵并过滤上三角(不含对角线) stacked_series = df.stack() filtered = stacked_series[stacked_series.index.get_level_values(0) < stacked_series.index.get_level_values(1)] # 转换为三列格式 result_df = filtered.reset_index() result_df.columns = ["item1", "item2", "distance"] print(result_df)
执行后输出符合要求的结果:
item1 item2 distance 0 1 2 1.154 1 1 3 1.235 2 1 4 1.297 3 1 5 0.960 4 2 3 0.932 5 2 4 0.929 6 2 5 0.988 7 3 4 0.727 8 3 5 1.244 9 4 5 1.019
2. 100,000×100,000规模矩阵的优化方案
10万级别的矩阵元素量达1e10,直接用pandas会内存爆炸,必须用numpy或分块处理:
方案A:numpy上三角索引提取
利用np.triu_indices_from()直接获取上三角(跳过对角线)的索引,内存效率远高于pandas:
import numpy as np # 假设dist_matrix是100000×100000的numpy对称矩阵 dist_matrix = np.random.rand(100000, 100000) dist_matrix = (dist_matrix + dist_matrix.T) / 2 # 构造对称矩阵 np.fill_diagonal(dist_matrix, 0) # 获取上三角(k=1表示跳过对角线)的行、列索引 rows, cols = np.triu_indices_from(dist_matrix, k=1) distances = dist_matrix[rows, cols] # 组合为三列数组(索引从1开始则加1) result_array = np.column_stack([rows + 1, cols + 1, distances])
注意:该方案需要约120GB内存(5e9个int64索引占80GB,5e9个float64距离占40GB),仅适用于大内存服务器。
方案B:分块处理(低内存场景)
将矩阵分块处理,边计算边写入磁盘,避免一次性加载所有数据:
import numpy as np import csv chunk_size = 1000 # 每次处理1000行,可根据内存调整 n = 100000 # 写入结果到CSV with open('distance_list.csv', 'w', newline='') as f: writer = csv.writer(f) writer.writerow(['item1', 'item2', 'distance']) for start_row in range(0, n, chunk_size): end_row = min(start_row + chunk_size, n) chunk = dist_matrix[start_row:end_row, :] for row_offset in range(chunk.shape[0]): global_row = start_row + row_offset # 只处理列索引大于当前行的部分 target_cols = range(global_row + 1, n) row_distances = chunk[row_offset, target_cols] # 逐行写入 for col, dist in zip(target_cols, row_distances): writer.writerow([global_row + 1, col + 1, dist])
该方案内存占用仅为chunk_size × n的矩阵内存(约800MB,float64),但处理时间较长,取决于磁盘IO速度。
3. 性能对比
- pandas stack():仅适用于≤1000×1000的矩阵,1000级规模耗时几秒;10万级规模直接内存溢出,完全不可用。
- numpy triu_indices:10万级规模在大内存服务器上耗时几十分钟,内存占用极高。
- 分块处理:10万级规模耗时数小时,内存占用低,适合普通硬件。
内容的提问来源于stack exchange,提问作者Philipp O.
相关产品推荐
相关产品推荐

