R语言手动实现Kernel K-Means嵌套循环运行过慢的优化求助
Hey, I see exactly where your bottleneck is—those nested loops calculating third are killing your performance, especially as your dataset grows. R is terrible at handling nested for loops with element-wise checks like data[k,'cluster']==g; those operations are slow because they’re interpreted line-by-line instead of using optimized compiled code. Let’s fix this with vectorized matrix operations, which are R’s sweet spot for numerical computations.
The Root of the Slowdown
Your original code for third runs a triple loop: for each cluster g, you iterate over every pair of samples (k,l) and accumulate the product of their cluster membership indicators and the kernel matrix value. That’s an O(c*n²) operation, which gets exponentially slow as n increases. Every element-wise check data[k,'cluster']==g adds unnecessary overhead, and nested loops in R simply aren’t built for this kind of heavy computation.
Optimized Solution for the third Calculation
Instead of checking membership for every pair, we can create an indicator matrix where each row corresponds to a sample, and each column corresponds to a cluster. The value indicator[i,g] is 1 if sample i is in cluster g, 0 otherwise. Then, third[g] is just the quadratic form of the kernel matrix with respect to the indicator vector for cluster g—something we can compute in one go with matrix multiplication (which uses optimized compiled code under the hood).
Here’s how to replace that entire slow loop block:
# Generate indicator matrix (n rows, c columns) n <- nrow(data) indicator <- matrix(0, nrow = n, ncol = c) indicator[cbind(1:n, data$cluster)] <- 1 # Set 1s where sample belongs to its cluster # Calculate cluster quadratic forms in one matrix operation cluster_kernel_quad <- t(indicator) %*% kernelmatrix %*% indicator third <- diag(cluster_kernel_quad) # Extract diagonal values for each cluster
This will cut your runtime for this step from minutes to seconds, even for large datasets.
Bonus: Speed Up the Distance Calculation Too
You can also eliminate the nested loops in the distance calculation by reusing the indicator matrix and precomputing cluster sizes (so you don’t recalculate sum(data[,'cluster']==j) every time):
# Precompute size of each cluster once per iteration cluster_sizes <- colSums(indicator) # Vectorize distance calculation (no more loops!) k_ii <- diag(kernelmatrix) # K(ii, ii) for all samples sum_k_ij <- kernelmatrix %*% indicator # Sum of K(ii, k) for k in cluster j, for all ii,j # Build the distance matrix in one vectorized step distance <- outer(k_ii, rep(1, c)) - 2 * sum_k_ij / cluster_sizes + third / (cluster_sizes^2)
Why This Works
R’s for loops are slow because they’re interpreted at runtime. Matrix and vector operations, however, leverage BLAS/LAPACK libraries—compiled code optimized for speed and memory efficiency. Even with 1000+ samples, this approach will drastically reduce your iteration time.
内容的提问来源于stack exchange,提问作者Mateus Maia

