You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言dplyr::arrange()对字母数字混合列(Q1/Q2等)排序失效问题

如何按Q1、Q2、Q3...的顺序排列rstatix统计结果的variable列?

我查阅过许多关于arrange()的相关帖子,但均未解决我的问题,希望此问题并非重复提问。我的数据包含名为Q1、Q2、Q3等的列,使用rstatix::get_summary_stats()计算基础描述统计量后,需要将新生成的variable列按升序排列(即Q1在Q2前,Q2在Q3前等)。我确信这是个简单问题,但找不到问题所在。

原始数据

ID Q1 Q2 Q3 Q4 Q5 Q6 Q7 Q8 Q9 Q10 Q11 Q12 Q13 Q14 Q15
1 PART1  4  1  1  5  5  5  1  5  1   1   3   5   5   1   5
2 PART2  5  4  1  5  5  4  1  5  2   1   3   5   4   1   5
3 PART3  2  4  3  5  5  4  1  5  2   1   3   5   4   1   5
so on...

尝试的代码

descriptive <-  data %>% 
  rstatix::get_summary_stats(show = c("mean", "sd", "median", "iqr", "min", "max"))  %>% 
  mutate_if(is.numeric, round, 2) %>% 
  dplyr::arrange(variable) 

当前结果(前10行)

A tibble: 15 x 8
   variable     n  mean    sd median   iqr   min   max
   <chr>    <dbl> <dbl> <dbl>  <dbl> <dbl> <dbl> <dbl>
 1 Q1          63  3.94  1.03      4   2       2     5
 2 Q10         63  1.84  0.88      2   2       1     3
 3 Q11         63  2.62  1.31      3   3       1     5
 4 Q12         63  3.98  1.01      4   2       2     5
 5 Q13         63  4.33  0.8       5   1       2     5
 6 Q14         63  1.91  0.88      2   2       1     4
 7 Q15         63  4.25  0.95      5   1       2     5
 8 Q2          63  2.86  1.58      3   3       1     5
 9 Q3          63  1.97  1.06      2   2       1     4
10 Q4          63  3.98  1.04      4   2       2     5

已尝试的无效操作

  • 使用ungroup()
  • 使用across(starts_with("Q*"))

数据结构

> dput(descriptive)[1:10, ]
structure(list(variable = c("Q1", "Q10", "Q11", "Q12", "Q13", 
"Q14", "Q15", "Q2", "Q3", "Q4", "Q5", "Q6", "Q7", "Q8", "Q9"), 
    n = c(63, 63, 63, 63, 63, 63, 63, 63, 63, 63, 63, 63, 63, 
    63, 63), mean = c(3.94, 1.84, 2.62, 3.98, 4.33, 1.91, 4.25, 
    2.86, 1.97, 3.98, 4.21, 4.05, 2.38, 4.03, 2.25), sd = c(1.03, 
    0.88, 1.31, 1.01, 0.8, 0.88, 0.95, 1.58, 1.06, 1.04, 0.94, 
    1.04, 1.36, 1.05, 1.12), median = c(4, 2, 3, 4, 5, 2, 5, 
    3, 2, 4, 4, 4, 2, 4, 2), iqr = c(2, 2, 3, 2, 1, 2, 1, 3, 
    2, 2, 1, 2, 2.5, 2, 2), min = c(2, 1, 1, 2, 2, 1, 2, 1, 1, 
    2, 2, 1, 1, 2, 1), max = c(5, 3, 5, 5, 5, 4, 5, 5, 4, 5, 
    5, 5, 5, 5, 5)), row.names = c(NA, -15L), class = c("tbl_df", 
"tbl", "data.frame"))

解决方案

问题核心是variable列是字符类型,默认字典序排序时,"Q10"会因为第二个字符"1"小于"Q2"的"2"而排在前面,而非按数字大小排序。以下两种方法可以解决:

方法1:提取变量名中的数字,按数值排序

借助stringr包提取变量名里的数字,转为整数后排序:

library(dplyr)
library(stringr)

descriptive <- data %>% 
  rstatix::get_summary_stats(show = c("mean", "sd", "median", "iqr", "min", "max")) %>% 
  mutate_if(is.numeric, round, 2) %>% 
  arrange(as.integer(str_extract(variable, "\\d+")))

方法2:将变量名转为有序因子,指定原始顺序

利用原始数据中Q列的顺序,把variable转为有序因子后排序:

library(dplyr)

# 获取原始数据中Q开头列的顺序
q_column_order <- colnames(data)[startsWith(colnames(data), "Q")]

descriptive <- data %>% 
  rstatix::get_summary_stats(show = c("mean", "sd", "median", "iqr", "min", "max")) %>% 
  mutate_if(is.numeric, round, 2) %>% 
  mutate(variable = factor(variable, levels = q_column_order, ordered = TRUE)) %>% 
  arrange(variable)

内容的提问来源于stack exchange,提问作者Larissa Cury

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:00:54