You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark 2.3 VectorAssembler报错问题咨询

Testing VectorAssembler with PySpark 2.3.0

我最近在使用PySpark 2.3.0测试VectorAssembler的功能,为了简化测试流程,我从一个规模更大的数据集中挑选了部分数值型(Double数据类型)列,创建了一个小型的测试DataFrame。

具体的操作代码如下:

# 定义需要提取的列列表
cols = ['index','host_listings_count','neighbourhood_group_cleansed',
'bathrooms','bedrooms','beds','square_feet', 'guests_included',
'review_scores_rating']

# 从原DataFrame中选取指定列生成测试用DataFrame
test = df[cols]

# 查看测试DataFrame的前3行数据
test.take(3)

执行上述代码后,得到的前3行结果(部分内容省略):

[Row(index=0, host_listings_count=1, neighbourh...]

这里额外提个小注意点:如果后续要使用VectorAssembler将这些列转换为特征向量,得留意neighbourhood_group_cleansed列——如果它是字符串类型,必须先通过StringIndexer或OneHotEncoder这类工具转换成数值型,否则VectorAssembler会抛出错误,因为它只支持数值类型的输入列。

内容的提问来源于stack exchange,提问作者Odisseo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:11:42