ggparcoord平行坐标图问题:数据显示不全与坐标轴顺序调整
问题描述
我用ggparcoord处理含152条观测的数据集,绘制平行坐标图,X轴指定为MCA标准化维度变量Std_Dim1至Std_Dim4,按cluster变量分组,遇到两个问题:
- 图中显示的点数远少于152,怀疑数据未全部展示,同时收到警告:
In summary.lm(lm(x ~ as.factor(classVar == class.names[i]))) : essentially perfect fit: summary may be unreliable,无法理解含义; - 图中
Std_Dim4排在X轴首位,希望按Std_Dim1到Std_Dim4的顺序排列坐标轴。
附数据片段及原代码:
# 数据片段 structure(list(ID = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20), cluster = structure(c(1L, 2L, 2L, 2L, 2L, 1L, 3L, 1L, 2L, 2L, 3L, 2L, 1L, 4L, 2L, 1L, 5L, 1L, 3L, 5L), levels = c("1", "2", "3", "4", "5", "6"), class = "factor"), Std_Dim1 = c(-0.0728923703380469, -0.339155028394344, 0.340473420392532, -0.339155028394344, -0.339155028394344, -0.0728923703380469, -0.267300038214234, -0.0728923703380469, -0.339155028394344, -0.339155028394344, -0.267300038214234, -0.339155028394344, -0.0728923703380469, 8.24372395142489, -0.339155028394344, -0.0728923703380469, -0.340751805953902, 1.23367086777029, -0.144946957713102, 0.895154025143995), Std_Dim2 = c(0.380193300283392, -0.21657689529506, -0.0605737618430261, -0.21657689529506, -0.21657689529506, 0.380193300283392, 0.0358222042004695, 0.380193300283392, -0.21657689529506, -0.21657689529506, 0.0358222042004695, -0.21657689529506, 0.380193300283392, 2.24175307931177, -0.21657689529506, 0.380193300283392, 4.41531912524421, 0.549235501606044, 0.127794200787863, 1.07941331484527), Std_Dim3 = c(-0.249302376419046, -0.301241388482618, -0.515922638345379, -0.301241388482618, -0.301241388482618, -0.249302376419046, -0.243613817954941, -0.249302376419046, -0.301241388482618, -0.301241388482618, -0.243613817954941, -0.301241388482618, -0.249302376419046, -2.55242656849512, -0.301241388482618, -0.249302376419046, 3.97061868943169, -0.505287507303791, -0.306929946946723, 0.594582905299551), Std_Dim4 = c(-1.78310815898572, 0.600501349348233, 0.970728782381305, 0.600501349348233, 0.600501349348233, -1.78310815898572, -0.595475416793659, -1.78310815898572, 0.600501349348233, 0.600501349348233, -0.595475416793659, 0.600501349348233, -1.78310815898572, 2.19760934092996, 0.600501349348233, -1.78310815898572, 0.0136383315437233, -1.23858333677849, -0.586822354919764, 1.81841980809893)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame"))
# 原代码 ggparcoord(res.cluster, columns = 21:24, groupColumn = 2, order = "anyClass", showPoints = TRUE, title = "Location of cluster points in the MCA dimensions", alphaLines = 0.5 ) + scale_color_viridis(discrete=TRUE) + theme_ipsum()+ theme( plot.title = element_text(size=10) )+ theme(text=element_text(family="Calibri"))+ xlab("")
解决方案
问题1:点数缺失与警告解释
警告含义
该警告来自ggparcoord内部的坐标轴排序逻辑:当设置order = "anyClass"时,函数会为每个类别拟合线性模型来确定坐标轴顺序。如果某个维度下,同一类别的观测值完全一致(比如数据片段中cluster 1的所有Std_Dim1值相同),线性模型会出现完全拟合,导致统计量失去意义,因此抛出该警告。
点数缺失原因
视觉上点数少是因为大量观测点重合——同一cluster的多条观测在所有维度上数值完全一致,绘图时重叠在一起,看起来像一个点。可以通过调整点的尺寸和透明度让重叠点更明显:
showPoints = TRUE, pointSize = 2, pointAlpha = 0.3
若不需要坐标轴自动排序,可将order设为"none",既能消除警告,也能保留原始变量顺序:
order = "none"
问题2:坐标轴顺序调整
核心解决方法
原代码用列索引21:24指定X轴变量,若数据集列顺序变动或索引对应错误,会导致坐标轴混乱。直接用变量名指定,同时配合order = "none"强制保留指定顺序:
columns = c("Std_Dim1", "Std_Dim2", "Std_Dim3", "Std_Dim4") order = "none"
修改后的完整代码
ggparcoord(res.cluster, columns = c("Std_Dim1", "Std_Dim2", "Std_Dim3", "Std_Dim4"), groupColumn = "cluster", order = "none", showPoints = TRUE, pointSize = 2, pointAlpha = 0.3, title = "Location of cluster points in the MCA dimensions", alphaLines = 0.5 ) + scale_color_viridis(discrete=TRUE) + theme_ipsum()+ theme( plot.title = element_text(size=10), text=element_text(family="Calibri") )+ xlab("")
内容的提问来源于stack exchange,提问作者elsich
相关产品推荐
相关产品推荐

