You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PySpark转换行列生成带对角线单位值的相似性数据框

PySpark生成对称相似矩阵实现方法

完整实现代码

from pyspark.sql.functions import lit

# 提取所有唯一短语,生成对角线1.0数据
all_phrases = df.select("phrase1").union(df.select("phrase2")).distinct().rdd.map(lambda x: x[0]).collect()
diag_data = [(phrase, phrase, 1.0) for phrase in all_phrases]
diag_df = spark.createDataFrame(diag_data, schema=["phrase1", "phrase2", "common_persent"])

# 生成反向配对数据,补全矩阵下三角
reverse_pair_df = df.selectExpr("phrase2 as phrase1", "phrase1 as phrase2", "common_persent")

# 合并原始数据、反向数据、对角线数据
full_pair_df = df.unionByName(reverse_pair_df).unionByName(diag_df)

# 透视生成最终相似矩阵
sim_matrix = full_pair_df.groupBy("phrase1")\
                         .pivot("phrase2")\
                         .agg({"common_persent": "first"})\
                         .withColumnRenamed("phrase1", "phrases")\
                         .orderBy("phrases")

# 输出结果
sim_matrix.show()

实现逻辑说明

  • 对角线填充:先提取所有唯一短语,手动生成每个短语与自身配对、相似度为1.0的数据集,直接解决对角线单位值的填充需求
  • 对称补全:原数据集只有单向的短语配对(仅包含上三角数据),将phrase1和phrase2字段互换生成反向配对数据,保证矩阵上下对称
  • 透视转换:合并所有配对数据后,通过groupBy + pivot的透视操作,将行转列得到最终的相似矩阵

内容的提问来源于stack exchange,提问作者Rory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 18:54:04