You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化大规模短语数据集上余弦相似度的运行效率

相似短语分组的性能优化问题

当前场景与问题

  • 目标:通过余弦相似度对相似短语进行分组,减少数据库中存储的短语条目数量
  • 实现基础:基于一篇Python方案实现,核心依赖sparse_dot_topn库(也是性能瓶颈所在)
  • 性能现状:处理1万篇文章、350万条短语时,余弦相似度计算耗时70分钟;计划处理100万篇文章,当前效率完全不可行,且运行时间随短语数量增长呈非线性上升

已尝试的优化方案及问题

  • 多线程版本sparse_dot_topn:使用基于Cython编写的多线程实现,虽有提速,但上述70分钟的耗时就是该方案的测试结果
  • 分批处理策略:每x篇文档执行一次分组,单批次时间缩短,但存在跨批次短语无法对比的问题;尝试将部分短语带入后续批次,又导致单批次运行时间随带入短语数量增加而变长
  • PySpark尝试:考虑用PySpark做分布式处理,但不清楚如何将自定义的awesome_cossim_top函数与PySpark集成

自定义匹配生成函数

# 计算两个TF-IDF向量的余弦相似度,结果为稀疏矩阵,Scikit-learn会返回CSR稀疏矩阵
def awesome_cossim_top(A, B, ntop, lower_bound=0):
    # 强制转换为CSR矩阵,若已是CSR则无额外开销
    A = A.tocsr()
    B = B.tocsr()
    M, _ = A.shape
    _, N = B.shape
 
    idx_dtype = np.int32
 
    nnz_max = M*ntop
 
    indptr = np.zeros(M+1, dtype=idx_dtype)
    indices = np.zeros(nnz_max, dtype=idx_dtype)
    data = np.zeros(nnz_max, dtype=A.dtype)
    n_jobs = 8

    ct.sparse_dot_topn_threaded(
        M, N, np.asarray(A.indptr, dtype=idx_dtype),
        np.asarray(A.indices, dtype=idx_dtype),
        A.data,
        np.asarray(B.indptr, dtype=idx_dtype),
        np.asarray(B.indices, dtype=idx_dtype),
        B.data,
        ntop,
        lower_bound,
        indptr, indices, data, n_jobs)

    return csr_matrix((data,indices,indptr),shape=(M,N))

内容的提问来源于stack exchange,提问作者ivan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 13:35:12