R语言中Kmeans预测工具包推荐及建模测试方案咨询
Hey Maria, sorry to hear the kmeans.predict() function from SwarmSVM is no longer available—let’s walk through how to split your dataframe into training/test sets and handle K-means clustering prediction in R, using both base R and some handy packages.
First, you’ll want to partition your data to train the model on a subset and assign clusters to unseen test data. Here are three reliable approaches:
Base R (No Extra Packages Needed)
Use the built-in sample() function to randomly select indices for your training set:
# Assume your dataframe is named `df` set.seed(123) # Fix random seed for reproducibility train_size <- floor(0.7 * nrow(df)) # 70% for training, 30% for test train_indices <- sample(seq_len(nrow(df)), size = train_size) train_df <- df[train_indices, ] test_df <- df[-train_indices, ]
Using the caret Package
The caret package has a dedicated function for stratified splits (great if you want to preserve proportions of a categorical variable in your splits):
library(caret) set.seed(123) # If you have a categorical column (e.g., `df$category`), use it for stratification train_indices <- createDataPartition(df$category, p = 0.7, list = FALSE) # If no categorical column, use `1:nrow(df)` instead # train_indices <- createDataPartition(1:nrow(df), p = 0.7, list = FALSE) train_df <- df[train_indices, ] test_df <- df[-train_indices, ]
Using the rsample Package
rsample makes splitting data clean and readable with its intuitive split functions:
library(rsample) set.seed(123) data_split <- initial_split(df, prop = 0.7) # 70% training, 30% test train_df <- training(data_split) test_df <- testing(data_split)
Base R’s kmeans() function doesn’t include a built-in predict() method, but we can either implement this ourselves or use a package that adds this functionality.
Base R Implementation (Custom Predict Function)
After training your K-means model on the training set, calculate distances from test samples to cluster centers and assign the closest cluster:
# Train K-means on the training set (e.g., 3 clusters) set.seed(123) kmeans_model <- kmeans(train_df, centers = 3) # Custom predict function for K-means predict_kmeans <- function(model, new_data) { # Calculate Euclidean distance from each test sample to every cluster center dist_matrix <- apply(new_data, 1, function(sample) { apply(model$centers, 1, function(center) { sqrt(sum((sample - center)^2)) }) }) # Assign each sample to the cluster with the smallest distance cluster_assignments <- apply(dist_matrix, 2, which.min) return(cluster_assignments) } # Assign clusters to test data test_cluster_labels <- predict_kmeans(kmeans_model, test_df)
Using the flexclust Package (Built-in Predict Method)
The flexclust package provides a kcca() function (equivalent to K-means) that includes a native predict() method, making this process much simpler:
library(flexclust) set.seed(123) # Train K-means model with kcca kcca_model <- kcca(train_df, k = 3, family = kccaFamily("kmeans")) # Directly predict clusters for test data test_cluster_labels <- predict(kcca_model, newdata = test_df)
- Always set a random seed (
set.seed()) to ensure your splits and clustering results are reproducible. - For K-means, make sure your data is scaled (using
scale(df)if needed) before training, since the algorithm is sensitive to feature scales.
内容的提问来源于stack exchange,提问作者Maria Gold

