在R中如何更简便地实现分类数据与多列数值数据的相关性分析
优化实现方案
首先安装所需依赖包,首次运行执行以下代码:
install.packages(c("dplyr", "readxl", "psych"))
核心运行代码如下,仅需修改文件路径即可直接运行:
library(dplyr) library(readxl) library(psych) # 读取完整Excel数据,替换为你的本地文件路径 df_full <- read_excel("你的调研数据文件.xlsx") # 自动对第一列分类数据做编码映射 col1_name <- colnames(df_full)[1] df_encoded <- df_full %>% mutate(!!sym(col1_name) := case_when( !!sym(col1_name) == "Seasoned Professional; with a variety of technical skills developed; strong problem solver" ~ 1, !!sym(col1_name) == "Entry-level; learning technical and professional skills" ~ 2, !!sym(col1_name) == "Early career; with basic skills development; work contributes to the success of team goals" ~ 3, !!sym(col1_name) == "Advanced; very senior; can solve most problems of issues that arise" ~ 4, !!sym(col1_name) == "Expert; has broad influence and experience within organization" ~ 5 )) # 提取第一列编码值与后续所有数值列 col1_data <- df_encoded[[1]] numeric_cols <- df_encoded[, -1] # 批量计算斯皮尔曼相关性 cor_res <- corr.test(col1_data, numeric_cols, method = "spearman", adjust = "none") # 整理为标准化分析表 cor_analysis_table <- data.frame( 数值列名 = colnames(numeric_cols), 相关系数 = round(as.numeric(cor_res$r), 4), p值 = round(as.numeric(cor_res$p), 4), 显著性 = case_when( as.numeric(cor_res$p) < 0.001 ~ "***", as.numeric(cor_res$p) < 0.01 ~ "**", as.numeric(cor_res$p) < 0.05 ~ "*", TRUE ~ "不显著" ) ) # 输出分析表 print(cor_analysis_table, row.names = FALSE) # 若需导出结果到Excel可运行以下代码 # install.packages("writexl") # library(writexl) # write_xlsx(cor_analysis_table, "相关性分析结果.xlsx")
代码说明
- 自动识别第一列作为分类字段完成编码,无需手动拆分文件/列
- 自动遍历后续所有数值列批量计算相关性,一次性生成完整分析表
- 分析表内置显著性标记,可直接用于报告撰写
- 若分类映射规则调整,仅需修改
case_when中的赋值逻辑即可
内容的提问来源于stack exchange,提问作者datatitanic
相关产品推荐
相关产品推荐

