内存高效生成带权边列表:基于特征DataFrame的余弦相似度计算
Got it, let's work through this problem step by step. Your goal is to generate an edge list with cosine similarity weights (excluding zero-weight edges) without building a full adjacency matrix—this is crucial for memory efficiency when dealing with large datasets, even if you have plenty of compute power to spare. Here's a practical, memory-friendly approach:
1. 先把类别特征转成数值向量
First off, your features are categorical (like "Red", "Green") but cosine similarity works with numerical vectors. We'll use one-hot encoding to convert these categories into 0/1 vectors, which preserves the categorical information without imposing any artificial ordering.
import pandas as pd from sklearn.preprocessing import OneHotEncoder # 加载你的数据(这里用示例数据) data = { 'ID': ['ID1', 'ID2', 'ID3', 'ID4'], 'Col1': ['Red', 'Yellow', 'Gray', 'Yellow'], 'Col2': ['Green', 'Green', 'Green', 'Green'], 'Col3': ['Blue', 'Orange', 'Yellow', 'Blue'] } df = pd.DataFrame(data).set_index('ID') # 初始化编码器,用稀疏输出更省内存(适合特征多的场景) encoder = OneHotEncoder(sparse_output=True) encoded_features = encoder.fit_transform(df) # 把ID和对应的向量索引映射起来,方便后续调用 id_index = df.index.tolist()
2. 高效计算余弦相似度,生成边列表
The big challenge here is avoiding the full O(n²) adjacency matrix. Instead, we'll only compute similarities for unique node pairs (i < j to avoid duplicates) and keep only pairs with non-zero similarity. Since you have plenty of compute power, we can use parallel processing to speed things up.
方法1:并行计算(适合中等至大规模数据)
This approach uses parallel loops to compute similarities without storing the full matrix:
from scipy.spatial.distance import cosine from joblib import Parallel, delayed import itertools # 生成所有不重复的ID对(只算i<j,减少一半计算量) id_pairs = list(itertools.combinations(id_index, 2)) # 定义计算单对相似度的函数 def compute_sim(pair): id1, id2 = pair # 从稀疏矩阵中获取对应向量 vec1 = encoded_features[id_index.index(id1)].toarray().flatten() vec2 = encoded_features[id_index.index(id2)].toarray().flatten() sim = 1 - cosine(vec1, vec2) # cosine distance = 1 - cosine similarity if sim > 0: # 过滤掉权重为0的边 return (id1, id2, round(sim, 2)) # 保留两位小数,可按需调整 # 用所有CPU核心并行计算 results = Parallel(n_jobs=-1)(delayed(compute_sim)(pair) for pair in id_pairs) # 过滤掉None值(即相似度为0的对) edge_list = [res for res in results if res is not None]
方法2:稀疏矩阵优化(适合超大规模数据)
If your dataset is extremely large, using sparse matrix operations can save even more memory. We'll compute the similarity matrix but keep it sparse, then extract only non-zero entries:
from sklearn.metrics.pairwise import cosine_similarity import scipy.sparse as sp # 计算余弦相似度,得到稀疏矩阵 sim_sparse = sp.csr_matrix(cosine_similarity(encoded_features)) # 获取所有非零相似度的位置和值 rows, cols, values = sp.find(sim_sparse) # 过滤掉重复对(i >= j)和零值(find已经过滤了零值) edge_list = [] for i, j, val in zip(rows, cols, values): if i < j and val > 0: edge_list.append((id_index[i], id_index[j], round(val, 2)))
3. 输出符合要求的格式
Once you have the edge list, you can format it exactly like your example:
# 转为DataFrame方便查看或保存 edge_df = pd.DataFrame(edge_list, columns=['ID1', 'ID2', 'Weight']) # 输出为空格分隔的文本格式 for row in edge_df.itertuples(index=False): print(f"{row.ID1} {row.ID2} {row.Weight}")
关键内存优化点
- Sparse matrices: Using sparse output for one-hot encoding and similarity matrices avoids storing huge amounts of zero values.
- Avoid duplicate pairs: Calculating only i < j cuts the number of computations in half and prevents duplicate edges in your list.
- Filter early: We discard zero-similarity pairs as soon as we compute them, so we never store unnecessary entries.
内容的提问来源于stack exchange,提问作者user6453877

