You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sparklyr类型错误:预测列应为VectorUDT而非DoubleType,求解决方案

Fixing Sparklyr "Wrong Type" Error for Classification Metrics

Hey there, let's sort out this type mismatch issue you're running into with sparklyr! That error message is telling us the prediction column needs to be a Spark Vector type (specifically org.apache.spark.mllib.linalg.VectorUDT) instead of a plain Double—and this is a super common gotcha when working with classification metrics like AUC.

Why This Happens

Spark's MLlib classification evaluators (like the one for AUC) expect probability/score vectors (e.g., for binary classification, a vector like [0.2, 0.8] representing the probability of each class) instead of direct category labels (0/1 as Doubles). If your model only outputs the raw prediction column with class labels, you'll hit this error.

Solution 1: Train the Model to Output Probability Vectors

The easiest fix is to make sure your classification model generates a probability column during training. For most sparklyr classification models (like logistic regression), you can specify this with the probability_col parameter:

# Retrain your model with probability column enabled
lr_model <- ml_logistic_regression(
  partition$train,
  formula = Survived ~ Pclass + Sex + Age + SibSp + Parch,  # Match your feature columns
  probability_col = "probability"  # This ensures a vector-type column is created
)

# Generate predictions—now you'll have both `prediction` (Double) and `probability` (Vector)
train_pred <- ml_predict(lr_model, partition$train)

Solution 2: Convert Existing Double Column to Vector Type

If you can't retrain the model, you can manually convert your existing prediction column into a vector. For binary classification, we'll create a vector where one element is the probability of the negative class and the other is the positive class:

# Define a UDF to convert Double values to Spark Vectors
double_to_binary_vector <- function(x) {
  # For binary classification: if x is 0, vector is [1, 0]; if x is 1, vector is [0, 1]
  ml_vector(c(1 - x, x))
}

# Register the UDF with Spark
sdf_register_udf(sc, "double_to_binary_vector", double_to_binary_vector, return_type = "vector")

# Apply the UDF to your prediction column
train_pred <- train_pred %>%
  mutate(probability = invoke("double_to_binary_vector", prediction))

Use the Vector Column for Metrics Calculation

Finally, when you calculate metrics like AUC, use the vector column (e.g., probability) instead of the prediction column:

# Calculate AUC using the probability vector column
auc_evaluator <- ml_binary_classification_evaluator(
  train_pred,
  label_col = "Survived",
  raw_prediction_col = "probability",
  metric_name = "areaUnderROC"
)

# Get the AUC score
auc_score <- ml_evaluate(auc_evaluator)
print(auc_score)

Quick Verification

Double-check that your train_pred dataframe now has a probability column with type vector—you can confirm this with:

glimpse(train_pred)

内容的提问来源于stack exchange,提问作者Solène Bernard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:18:40