sparklyr类型错误:预测列应为VectorUDT而非DoubleType,求解决方案
Hey there, let's sort out this type mismatch issue you're running into with sparklyr! That error message is telling us the prediction column needs to be a Spark Vector type (specifically org.apache.spark.mllib.linalg.VectorUDT) instead of a plain Double—and this is a super common gotcha when working with classification metrics like AUC.
Why This Happens
Spark's MLlib classification evaluators (like the one for AUC) expect probability/score vectors (e.g., for binary classification, a vector like [0.2, 0.8] representing the probability of each class) instead of direct category labels (0/1 as Doubles). If your model only outputs the raw prediction column with class labels, you'll hit this error.
Solution 1: Train the Model to Output Probability Vectors
The easiest fix is to make sure your classification model generates a probability column during training. For most sparklyr classification models (like logistic regression), you can specify this with the probability_col parameter:
# Retrain your model with probability column enabled lr_model <- ml_logistic_regression( partition$train, formula = Survived ~ Pclass + Sex + Age + SibSp + Parch, # Match your feature columns probability_col = "probability" # This ensures a vector-type column is created ) # Generate predictions—now you'll have both `prediction` (Double) and `probability` (Vector) train_pred <- ml_predict(lr_model, partition$train)
Solution 2: Convert Existing Double Column to Vector Type
If you can't retrain the model, you can manually convert your existing prediction column into a vector. For binary classification, we'll create a vector where one element is the probability of the negative class and the other is the positive class:
# Define a UDF to convert Double values to Spark Vectors double_to_binary_vector <- function(x) { # For binary classification: if x is 0, vector is [1, 0]; if x is 1, vector is [0, 1] ml_vector(c(1 - x, x)) } # Register the UDF with Spark sdf_register_udf(sc, "double_to_binary_vector", double_to_binary_vector, return_type = "vector") # Apply the UDF to your prediction column train_pred <- train_pred %>% mutate(probability = invoke("double_to_binary_vector", prediction))
Use the Vector Column for Metrics Calculation
Finally, when you calculate metrics like AUC, use the vector column (e.g., probability) instead of the prediction column:
# Calculate AUC using the probability vector column auc_evaluator <- ml_binary_classification_evaluator( train_pred, label_col = "Survived", raw_prediction_col = "probability", metric_name = "areaUnderROC" ) # Get the AUC score auc_score <- ml_evaluate(auc_evaluator) print(auc_score)
Quick Verification
Double-check that your train_pred dataframe now has a probability column with type vector—you can confirm this with:
glimpse(train_pred)
内容的提问来源于stack exchange,提问作者Solène Bernard

