You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

内存高效生成带权边列表:基于特征DataFrame的余弦相似度计算

Got it, let's work through this problem step by step. Your goal is to generate an edge list with cosine similarity weights (excluding zero-weight edges) without building a full adjacency matrix—this is crucial for memory efficiency when dealing with large datasets, even if you have plenty of compute power to spare. Here's a practical, memory-friendly approach:

解决方案步骤

1. 先把类别特征转成数值向量

First off, your features are categorical (like "Red", "Green") but cosine similarity works with numerical vectors. We'll use one-hot encoding to convert these categories into 0/1 vectors, which preserves the categorical information without imposing any artificial ordering.

import pandas as pd
from sklearn.preprocessing import OneHotEncoder

# 加载你的数据(这里用示例数据)
data = {
    'ID': ['ID1', 'ID2', 'ID3', 'ID4'],
    'Col1': ['Red', 'Yellow', 'Gray', 'Yellow'],
    'Col2': ['Green', 'Green', 'Green', 'Green'],
    'Col3': ['Blue', 'Orange', 'Yellow', 'Blue']
}
df = pd.DataFrame(data).set_index('ID')

# 初始化编码器,用稀疏输出更省内存(适合特征多的场景)
encoder = OneHotEncoder(sparse_output=True)
encoded_features = encoder.fit_transform(df)

# 把ID和对应的向量索引映射起来,方便后续调用
id_index = df.index.tolist()

2. 高效计算余弦相似度,生成边列表

The big challenge here is avoiding the full O(n²) adjacency matrix. Instead, we'll only compute similarities for unique node pairs (i < j to avoid duplicates) and keep only pairs with non-zero similarity. Since you have plenty of compute power, we can use parallel processing to speed things up.

方法1:并行计算(适合中等至大规模数据)

This approach uses parallel loops to compute similarities without storing the full matrix:

from scipy.spatial.distance import cosine
from joblib import Parallel, delayed
import itertools

# 生成所有不重复的ID对(只算i<j,减少一半计算量)
id_pairs = list(itertools.combinations(id_index, 2))

# 定义计算单对相似度的函数
def compute_sim(pair):
    id1, id2 = pair
    # 从稀疏矩阵中获取对应向量
    vec1 = encoded_features[id_index.index(id1)].toarray().flatten()
    vec2 = encoded_features[id_index.index(id2)].toarray().flatten()
    sim = 1 - cosine(vec1, vec2)  # cosine distance = 1 - cosine similarity
    if sim > 0:  # 过滤掉权重为0的边
        return (id1, id2, round(sim, 2))  # 保留两位小数,可按需调整

# 用所有CPU核心并行计算
results = Parallel(n_jobs=-1)(delayed(compute_sim)(pair) for pair in id_pairs)

# 过滤掉None值(即相似度为0的对)
edge_list = [res for res in results if res is not None]

方法2:稀疏矩阵优化(适合超大规模数据)

If your dataset is extremely large, using sparse matrix operations can save even more memory. We'll compute the similarity matrix but keep it sparse, then extract only non-zero entries:

from sklearn.metrics.pairwise import cosine_similarity
import scipy.sparse as sp

# 计算余弦相似度,得到稀疏矩阵
sim_sparse = sp.csr_matrix(cosine_similarity(encoded_features))

# 获取所有非零相似度的位置和值
rows, cols, values = sp.find(sim_sparse)

# 过滤掉重复对(i >= j)和零值(find已经过滤了零值)
edge_list = []
for i, j, val in zip(rows, cols, values):
    if i < j and val > 0:
        edge_list.append((id_index[i], id_index[j], round(val, 2)))

3. 输出符合要求的格式

Once you have the edge list, you can format it exactly like your example:

# 转为DataFrame方便查看或保存
edge_df = pd.DataFrame(edge_list, columns=['ID1', 'ID2', 'Weight'])

# 输出为空格分隔的文本格式
for row in edge_df.itertuples(index=False):
    print(f"{row.ID1} {row.ID2} {row.Weight}")

关键内存优化点

  • Sparse matrices: Using sparse output for one-hot encoding and similarity matrices avoids storing huge amounts of zero values.
  • Avoid duplicate pairs: Calculating only i < j cuts the number of computations in half and prevents duplicate edges in your list.
  • Filter early: We discard zero-similarity pairs as soon as we compute them, so we never store unnecessary entries.

内容的提问来源于stack exchange,提问作者user6453877

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:37:05