You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中Kmeans预测工具包推荐及建模测试方案咨询

Hey Maria, sorry to hear the kmeans.predict() function from SwarmSVM is no longer available—let’s walk through how to split your dataframe into training/test sets and handle K-means clustering prediction in R, using both base R and some handy packages.

1. Splitting Your Dataframe into Training & Test Sets

First, you’ll want to partition your data to train the model on a subset and assign clusters to unseen test data. Here are three reliable approaches:

Base R (No Extra Packages Needed)

Use the built-in sample() function to randomly select indices for your training set:

# Assume your dataframe is named `df`
set.seed(123) # Fix random seed for reproducibility
train_size <- floor(0.7 * nrow(df)) # 70% for training, 30% for test
train_indices <- sample(seq_len(nrow(df)), size = train_size)

train_df <- df[train_indices, ]
test_df <- df[-train_indices, ]

Using the caret Package

The caret package has a dedicated function for stratified splits (great if you want to preserve proportions of a categorical variable in your splits):

library(caret)
set.seed(123)

# If you have a categorical column (e.g., `df$category`), use it for stratification
train_indices <- createDataPartition(df$category, p = 0.7, list = FALSE)
# If no categorical column, use `1:nrow(df)` instead
# train_indices <- createDataPartition(1:nrow(df), p = 0.7, list = FALSE)

train_df <- df[train_indices, ]
test_df <- df[-train_indices, ]

Using the rsample Package

rsample makes splitting data clean and readable with its intuitive split functions:

library(rsample)
set.seed(123)

data_split <- initial_split(df, prop = 0.7) # 70% training, 30% test
train_df <- training(data_split)
test_df <- testing(data_split)
2. Training K-Means & Assigning Clusters to Test Data

Base R’s kmeans() function doesn’t include a built-in predict() method, but we can either implement this ourselves or use a package that adds this functionality.

Base R Implementation (Custom Predict Function)

After training your K-means model on the training set, calculate distances from test samples to cluster centers and assign the closest cluster:

# Train K-means on the training set (e.g., 3 clusters)
set.seed(123)
kmeans_model <- kmeans(train_df, centers = 3)

# Custom predict function for K-means
predict_kmeans <- function(model, new_data) {
  # Calculate Euclidean distance from each test sample to every cluster center
  dist_matrix <- apply(new_data, 1, function(sample) {
    apply(model$centers, 1, function(center) {
      sqrt(sum((sample - center)^2))
    })
  })
  # Assign each sample to the cluster with the smallest distance
  cluster_assignments <- apply(dist_matrix, 2, which.min)
  return(cluster_assignments)
}

# Assign clusters to test data
test_cluster_labels <- predict_kmeans(kmeans_model, test_df)

Using the flexclust Package (Built-in Predict Method)

The flexclust package provides a kcca() function (equivalent to K-means) that includes a native predict() method, making this process much simpler:

library(flexclust)
set.seed(123)

# Train K-means model with kcca
kcca_model <- kcca(train_df, k = 3, family = kccaFamily("kmeans"))

# Directly predict clusters for test data
test_cluster_labels <- predict(kcca_model, newdata = test_df)
Quick Notes
  • Always set a random seed (set.seed()) to ensure your splits and clustering results are reproducible.
  • For K-means, make sure your data is scaled (using scale(df) if needed) before training, since the algorithm is sensitive to feature scales.

内容的提问来源于stack exchange,提问作者Maria Gold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:29:36