featurePlot函数输出NULL问题排查:功能限制或代码错误?
我使用Kaggle学生就业数据集开展机器学习实验,验证类别是否可由其他预测变量判定。尝试用caret包的featurePlot函数绘制各特征与类别的箱线图和密度图,但运行代码后输出NULL。请问该函数是否存在使用限制(比如9个1列的图数量过多),或是我的代码存在错误?
原代码:
dataset <- Student_Employability_Datasets #creating a reproducible dataset dataset_sub <- dput(dataset[1:20,]) dataset_struct <- structure( list( `Name of Student` = c( "Student 1", "Student 2", "Student 3", "Student 4", "Student 5", "Student 6", "Student 7", "Student 8", "Student 9", "Student 10", "Student 11", "Student 12", "Student 13", "Student 14", "Student 15", "Student 16", "Student 17", "Student 18", "Student 19", "Student 20" ), `General Appearance` = c(4, 4, 4, 3, 4, 4, 4, 5, 4, 4, 5, 3, 4, 3, 4, 3, 5, 4, 4, 4), `Speaking Manner` = c(5, 4, 3, 3, 4, 4, 4, 3, 4, 4, 5, 4, 3, 3, 4, 3, 3, 4, 4, 3), `Physical Condition` = c(4, 4, 3, 3, 3, 3, 4, 3, 4, 3, 5, 4, 3, 3, 3, 3, 3, 4, 4, 3), `Mental Alertness` = c(5, 4, 3, 2, 3, 3, 3, 4, 4, 4, 5, 4, 2, 3, 4, 3, 3, 4, 5, 4), `Self-Confidence` = c(5, 4, 3, 3, 4, 3, 3, 3, 4, 3, 5, 3, 3, 3, 4, 3, 3, 4, 5, 5), `Presenting Ideas` = c(5, 4, 3, 3, 4, 3, 3, 3, 4, 4, 5, 4, 3, 2, 4, 3, 3, 4, 4, 4), `Communication Skills` = c(5, 3, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 2, 2, 3, 3, 3, 4, 4, 3), `Student Performance Rating` = c(5, 5, 5, 5, 5, 5, 3, 5, 5, 5, 5, 5, 5, 4, 5, 4, 4, 5, 5, 5), Class = c(1, 1, 0, 0, 1, 1, 1, 1, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0) ), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame") ) dataset <- dataset_struct # changing some values to numerics dataset <- dataset %>% mutate_at(c('GENERAL APPEARANCE', 'MANNER OF SPEAKING', 'PHYSICAL CONDITION', 'MENTAL ALERTNESS', 'SELF-CONFIDENCE', 'ABILITY TO PRESENT IDEAS', 'COMMUNICATION SKILLS', 'Student Performance Rating'), as.numeric) sapply(dataset, class) # changing the names of the column titles names(dataset)[10] <- "Class" # changing the class column to 0's and 1's dataset$Class <- ifelse(dataset$Class == "Employable", 1, 0) # creating a list of 80% of the rows in the original dataset for training validation_index2 <- createDataPartition(dataset$Class, p=0.8, list = FALSE) # selecting the 20% of the data for validation validation <- dataset[-validation_index2, ] # using the remaining 80% of the data to train and test the model dataset <- dataset[validation_index2, ] # spliting input and output x <- dataset[, 1:9] y <- dataset[, 10] #box and whisker plots for each attribute(ERROR) featurePlot(x=x, y=y, plot = "box") #density plots for each attribute by class value (ERROR) scales <- list(x=list(relation="free"), y=list(relation="free")) featurePlot(x=x , y=y, plot="density", scales = scales)
错误分析
输出NULL的核心原因是代码中的多处逻辑错误,和featurePlot的图数量限制无关:
列名不匹配导致数值转换失效
代码中mutate_at使用的列名(如GENERAL APPEARANCE、MANNER OF SPEAKING)和dataset_struct中定义的实际列名(如General Appearance、Speaking Manner)完全不一致,导致数值类型转换没有生效。不过dataset_struct中的特征列本身已经是数值型,这一步其实多余,但列名不匹配会引发后续潜在问题。包含非数值型特征列
拆分输入输出时,x <- dataset[, 1:9]包含了第一列Name of Student(字符型),而featurePlot要求输入的x必须是数值型的预测变量矩阵,字符列会干扰绘图逻辑。Class列处理错误产生NA
dataset_struct中的Class列已经是0/1数值,但后续执行dataset$Class <- ifelse(dataset$Class == "Employable", 1, 0)时,数值和字符串比较会返回NA,导致y中出现缺失值,最终使featurePlot无法正常绘图,输出NULL。
修正后的代码
library(caret) library(dplyr) # 使用提供的可复现数据集 dataset_struct <- structure( list( `Name of Student` = c( "Student 1", "Student 2", "Student 3", "Student 4", "Student 5", "Student 6", "Student 7", "Student 8", "Student 9", "Student 10", "Student 11", "Student 12", "Student 13", "Student 14", "Student 15", "Student 16", "Student 17", "Student 18", "Student 19", "Student 20" ), `General Appearance` = c(4, 4, 4, 3, 4, 4, 4, 5, 4, 4, 5, 3, 4, 3, 4, 3, 5, 4, 4, 4), `Speaking Manner` = c(5, 4, 3, 3, 4, 4, 4, 3, 4, 4, 5, 4, 3, 3, 4, 3, 3, 4, 4, 3), `Physical Condition` = c(4, 4, 3, 3, 3, 3, 4, 3, 4, 3, 5, 4, 3, 3, 3, 3, 3, 4, 4, 3), `Mental Alertness` = c(5, 4, 3, 2, 3, 3, 3, 4, 4, 4, 5, 4, 2, 3, 4, 3, 3, 4, 5, 4), `Self-Confidence` = c(5, 4, 3, 3, 4, 3, 3, 3, 4, 3, 5, 3, 3, 3, 4, 3, 3, 4, 5, 5), `Presenting Ideas` = c(5, 4, 3, 3, 4, 3, 3, 3, 4, 4, 5, 4, 3, 2, 4, 3, 3, 4, 4, 4), `Communication Skills` = c(5, 3, 2, 3, 3, 3, 3, 3, 4, 4, 4, 4, 2, 2, 3, 3, 3, 4, 4, 3), `Student Performance Rating` = c(5, 5, 5, 5, 5, 5, 3, 5, 5, 5, 5, 5, 5, 4, 5, 4, 4, 5, 5, 5), Class = c(1, 1, 0, 0, 1, 1, 1, 1, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0) ), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame") ) dataset <- dataset_struct # 移除无关的学生名字列,保留数值特征和类别列 dataset <- dataset %>% select(-`Name of Student`) # 确认Class列类型(转换为因子型,符合featurePlot对分类变量的要求) dataset$Class <- as.factor(dataset$Class) # 创建训练集索引 validation_index2 <- createDataPartition(dataset$Class, p=0.8, list = FALSE) validation <- dataset[-validation_index2, ] dataset <- dataset[validation_index2, ] # 拆分输入输出:x取所有数值特征,y取类别列 x <- dataset %>% select(-Class) y <- dataset$Class # 绘制箱线图 featurePlot(x=x, y=y, plot = "box") # 绘制密度图 scales <- list(x=list(relation="free"), y=list(relation="free")) featurePlot(x=x , y=y, plot="density", scales = scales)
关键修正点说明
- 移除了
Name of Student字符列,确保x仅包含数值型特征。 - 将
Class转换为因子型:featurePlot要求目标变量y为因子(分类变量),这样绘图时会按类别区分展示。 - 删除了无效的列名匹配转换步骤,以及错误的Class列重新赋值逻辑。
- 使用
dplyr::select更清晰地选择特征列,避免索引带来的错误。
内容的提问来源于stack exchange,提问作者Parvitha Ramesh Rao

