You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark中使用RowMatrix计算columnSimilarities时类型转换错误求助

问题解决:TypeError: Cannot convert type <class 'pyspark.sql.types.Row'> into Vector

你的错误原因很明确:RowMatrix要求输入的RDD元素必须是Vector类型,但你直接传入的是包含Row对象的RDD——每个Row里才封装着features向量,所以Spark无法自动完成类型转换。

修复代码

只需要先从Row中提取出features字段,将RDD转换成由Vector组成的RDD即可:

# 提取Row中的features向量,得到Vector类型的RDD
vector_rdd = df.rdd.map(lambda row: row.features)
# 创建符合要求的RowMatrix
mat = RowMatrix(vector_rdd)
# 计算列相似度
sims = mat.columnSimilarities()

关键说明

  • 原代码里df.rdd返回的是RDD[Row],每个元素是包含features字段的Row对象,完全不符合RowMatrix的输入规范
  • 通过map(lambda row: row.features)可以直接取出每个Row内的DenseVector,得到RDD[Vector],这才是RowMatrix需要的标准输入格式

内容的提问来源于stack exchange,提问作者Johnas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:30:59