R语言kamila聚类包报错:Input datasets must be dataframes(单连续变量)
Hey, I’ve run into this exact quirk with kamila before! Let me break down what’s happening and how to fix it quickly:
When you pull just one continuous variable using data[, conInd], R automatically simplifies the result to a vector instead of keeping it as a data frame. But the kamila() function strictly expects the continuous variable input to be a data frame, which is why you hit that error. When you use two variables, the extraction stays as a data frame by default, so everything works smoothly.
Here’s the straightforward fix:
Add the drop=FALSE parameter when extracting your single continuous variable—this forces R to preserve the data frame structure instead of collapsing it to a vector:
library(kamila) data <- read.csv("mixed.csv", header=FALSE, sep=";") conInd <- 9 # This line is the key change: drop=FALSE keeps the result as a data frame conVars <- data[, conInd, drop=FALSE] conVars <- data.frame(scale(conVars)) catVarsFac <- data[, c(1,2,3,4,5,6,7,8)] catVarsFac[] <- lapply(catVarsFac, factor) kamRes <- kamila(conVars, catVarsFac, numClust=5, numInit=10, calcNumClust = "ps", numPredStrCvRun = 10, predStrThresh = 0.5)
Quick check to confirm it works:
Before running kamila(), you can verify conVars is a data frame with this quick base R command:
is.data.frame(conVars)
It should return TRUE after making the change.
If you’re comfortable with the tidyverse, another option is using dplyr::select() to pull the column—it always returns a data frame even for single columns—but the drop=FALSE trick is simpler if you’re sticking to base R.
内容的提问来源于stack exchange,提问作者Emrah BILGIC

