使用mutate(across)时fct_reorder结合str_replace排序错误排查
问题与解决方案:用
mutate(across)实现按对应_num列排序生成因子列 原始数据框
df <- tibble::tribble( ~how_bright_txt, ~how_bright_num, ~how_hard_txt, ~how_hard_num, ~how_hot_txt, ~how_hot_num, "Not very much", 1L, "Really hard", 5L, "Cold", 1L, "Somewhat", 2L, "Somewhat hard", 2L, "A bit cold", 2L, "Medium", 3L, "Medium", 3L, "Medium", 3L, "Quite a bit", 4L, "Quite hard", 4L, "Quite hot", 4L, "A lot", 5L, "Not very hard", 1L, "Really hot", 5L )
需求说明
生成新列:
- 列名为原列名去除
_txt或_num后缀 - 值取自对应的
_txt列 - 按对应
_num列的数值排序转换为因子
可行但重复的实现代码
通过多次调用fct_reorder可以实现需求,代码如下:
require(tidyverse) df %>% mutate(how_bright = fct_reorder(how_bright_txt, -how_bright_num), how_hard = fct_reorder(how_hard_txt, -how_hard_num), how_hot = fct_reorder(how_hot_txt, -how_hot_num)) %>% select(-c(ends_with("_txt"), ends_with("_num")))
尝试简化但失效的代码
尝试用mutate(across)简化代码,但生成的因子列排序不符合预期,与原_num列的排序不匹配:
df %>% mutate(across(ends_with("_txt"), ~ fct_reorder(.x, str_replace(.x, "_txt", "_num")), .names = '{stringr::str_remove({col}, "_txt")}')) %>% select(-c(ends_with("_txt"), ends_with("_num")))
问题原因分析
核心错误在于str_replace(.x, "_txt", "_num")的用法:
.x在这里是当前_txt列的向量值(比如how_bright_txt列的元素是"Not very much"、"Somewhat"这类字符串),而不是列名- 这些字符串里根本不含
_txt子串,所以str_replace返回的还是原字符串向量 fct_reorder会基于字符串的字典序排序,自然和_num列的数值排序逻辑不匹配
正确的简化实现
要获取对应_num列的数值,需要先拿到当前处理的列名,再转换为对应的数值列名,最后引用该列。代码如下:
df %>% mutate(across(ends_with("_txt"), ~ fct_reorder(.x, -.data[[str_replace(cur_column(), "_txt", "_num")]]), .names = '{str_remove(col, "_txt")}')) %>% select(-c(ends_with("_txt"), ends_with("_num")))
关键逻辑说明:
cur_column():获取当前正在处理的_txt列的列名(如how_bright_txt)str_replace(cur_column(), "_txt", "_num"):将列名转换为对应的_num列名(如how_bright_num).data[[...]]:引用转换后的数值列,加上负号实现降序排序,和原重复调用代码的逻辑完全一致
内容的提问来源于stack exchange,提问作者C.Robin
相关产品推荐
相关产品推荐

