如何为基于Base R实现的Data.frame排序函数添加按列单独设置排序方向的功能
给Base R实现的data.frame排序函数添加列级排序方向控制
我最近在写一个只靠Base R实现的data.frame排序函数sortdf,专门给非交互式场景用,而且只接受data.frame类型的输入。目前这个函数已经能按单列或多列做全局升序/降序排序,但现在想加个实用功能:给by参数指定的每一列单独设置排序方向,比如像这样调用:
sortdf(iris,by=c("Sepal.Width","Petal.Width"),dir=c("up","down"))
先给大家看看我现在的函数实现:
#' @title Sort a data.frame #' @description Sort a data.frame based on one or more columns #' @param x A data.frame class object #' @param by A column in the data.frame. Defaults to NULL, which sorts by all columns. #' @param decreasing A logical indicating the direction of sorting. #' @return A data.frame. #' sortdf <- function(x,by=NULL,decreasing=FALSE) { if(!is.data.frame(x)) stop("Input is not a data.frame.") if(is.null(by)) { ord <- do.call(order,x) } else { if(any(!by %in% colnames(x))) stop("One or more items in 'by' was not found.") if(length(by) == 1) ord <- order(x[ , by]) if(length(by) > 1) ord <- do.call(order, x[ , by]) } if(decreasing) ord <- rev(ord) return(x[ord, , drop=FALSE]) }
现有的几种调用方式:
sortdf(iris) # 按所有列升序 sortdf(iris,"Petal.Length") # 按Petal.Length升序 sortdf(iris,"Petal.Length",decreasing=TRUE) # 按Petal.Length降序 sortdf(iris,c("Petal.Length","Sepal.Length")) # 按两列升序
接下来分享几个可行的实现思路,都是纯Base R的方案:
思路1:新增dir参数,支持友好的字符或逻辑向量输入
我个人比较推荐这个方案,因为dir参数的语义更清晰,而且支持用户用更直观的"up"/"down"来指定方向,同时也兼容逻辑值。
具体实现步骤分这几步:
- 参数校验与转换:
- 先把字符型的
dir转换成逻辑值:"down"对应TRUE(降序),"up"对应FALSE(升序) - 确保
dir的长度和by匹配:如果用户只传了单个方向,就自动扩展到和by列数一致;如果by是NULL(按所有列排序),就扩展到所有列的数量 - 还要检查
by里的列名是否都存在于输入的data.frame中
- 先把字符型的
- 构造排序列列表:
- 对每个要排序的列,根据
dir的值决定是否取反。这里要用xtfrm()函数来处理所有可排序的类型(比如因子、字符、数值),因为直接用-对因子列会报错,xtfrm会把任意可排序对象转换成对应的数值,取反后就能实现降序
- 对每个要排序的列,根据
- 生成排序索引并返回结果:
- 用
do.call(order, ...)把处理后的列列表传进去,得到排序索引,再用这个索引提取data.frame的行
- 用
修改后的函数代码如下:
#' @title Sort a data.frame with per-column direction control #' @description Sort a data.frame based on one or more columns, allowing separate sort directions for each column #' @param x A data.frame class object #' @param by Character vector of column names to sort by. Defaults to NULL (sort by all columns). #' @param dir Vector specifying sort direction for each column in `by`. Can be logical (TRUE = decreasing, FALSE = increasing) or character ("up" = increasing, "down" = decreasing). Defaults to FALSE (all columns sorted in increasing order). #' @return A sorted data.frame. #' sortdf <- function(x, by = NULL, dir = FALSE) { # 校验输入是否为data.frame if (!is.data.frame(x)) stop("Input must be a data.frame.") # 处理dir参数,统一转换成逻辑向量(TRUE=降序,FALSE=升序) if (is.character(dir)) { dir <- match.arg(dir, choices = c("up", "down"), several.ok = TRUE) dir <- dir == "down" } if (!is.logical(dir)) stop("dir must be a logical or character vector.") # 处理by为NULL的情况:默认按所有列排序 if (is.null(by)) { by <- colnames(x) # 如果只传了一个方向,扩展到所有列 if (length(dir) == 1) dir <- rep(dir, length(by)) } else { # 校验指定的列是否存在 missing_cols <- setdiff(by, colnames(x)) if (length(missing_cols) > 0) { stop(sprintf("Columns not found: %s", paste(missing_cols, collapse = ", "))) } # 统一dir的长度:单个值扩展到和by一致 if (length(dir) == 1) dir <- rep(dir, length(by)) # 校验dir和by长度匹配 if (length(dir) != length(by)) { stop("Length of dir must match length of by.") } } # 构造用于排序的列列表:对降序列取反(用xtfrm实现通用兼容) sort_columns <- lapply(seq_along(by), function(i) { col_data <- x[[by[i]]] if (dir[i]) -xtfrm(col_data) else col_data }) # 生成排序索引 sort_order <- do.call(order, sort_columns) # 返回排序后的data.frame,drop=FALSE避免变成向量 return(x[sort_order, , drop = FALSE]) }
测试一下这个函数:
# 按Sepal.Width升序,Petal.Width降序 sortdf(iris, by = c("Sepal.Width", "Petal.Width"), dir = c("up", "down")) # 用逻辑向量指定方向,效果一样 sortdf(iris, by = c("Sepal.Width", "Petal.Width"), dir = c(FALSE, TRUE)) # 兼容原有调用方式,全局降序 sortdf(iris, by = "Petal.Length", dir = "down") # 按所有列降序 sortdf(iris, dir = "down")
思路2:改造原有decreasing参数,支持向量输入
如果你不想新增参数,也可以直接改造现有的decreasing参数,让它支持向量输入,这样更贴合Base R的习惯(虽然原生order的decreasing是单个值,但我们可以自己处理多列的情况)。
改造后的函数大概是这样:
#' @title Sort a data.frame with per-column direction control #' @description Sort a data.frame based on one or more columns, allowing separate sort directions for each column via the decreasing parameter #' @param x A data.frame class object #' @param by Character vector of column names to sort by. Defaults to NULL (sort by all columns). #' @param decreasing Logical vector specifying sort direction for each column in `by`. TRUE = decreasing, FALSE = increasing. Can be a single value to apply to all columns. Defaults to FALSE. #' @return A sorted data.frame. #' sortdf <- function(x, by = NULL, decreasing = FALSE) { if (!is.data.frame(x)) stop("Input must be a data.frame.") # 处理by为NULL的情况 if (is.null(by)) { by <- colnames(x) } else { # 校验列是否存在 missing_cols <- setdiff(by, colnames(x)) if (length(missing_cols) > 0) { stop(sprintf("Columns not found: %s", paste(missing_cols, collapse = ", "))) } } # 统一decreasing的长度:单个值扩展到和by一致 if (length(decreasing) == 1) { decreasing <- rep(decreasing, length(by)) } # 校验长度匹配 if (length(decreasing) != length(by)) { stop("Length of decreasing must match length of by.") } # 构造排序列列表 sort_columns <- lapply(seq_along(by), function(i) { col_data <- x[[by[i]]] if (decreasing[i]) -xtfrm(col_data) else col_data }) sort_order <- do.call(order, sort_columns) x[sort_order, , drop = FALSE] }
调用方式就变成:
# 按Sepal.Width升序,Petal.Width降序 sortdf(iris, by = c("Sepal.Width", "Petal.Width"), decreasing = c(FALSE, TRUE))
几个关键注意事项
- 通用类型兼容:一定要用
xtfrm()来处理列,不要直接用-,不然因子列会报错。xtfrm是Base R专门用来把可排序对象转换成数值的函数,对因子、字符、日期等类型都有效。 - 向后兼容性:不管用哪种方案,都要确保原来的调用方式依然能正常工作,比如
sortdf(iris, "Petal.Length", decreasing = TRUE)不能失效。 - 参数校验要严谨:要检查列名是否存在、方向参数的长度是否匹配,这样用户调用时能得到清晰的错误提示,而不是莫名其妙的运行时错误。
- 性能表现:这种用
lapply构造列列表再do.call(order)的方式,性能和原生Base R的order调用几乎一样,从你提供的基准测试代码来看,肯定比dplyr的arrange快很多,符合非交互式场景的性能需求。
基准测试验证
你可以把修改后的sortdf加入到你的基准测试代码里,对比原生Base R和dplyr的性能:
library(microbenchmark) library(ggplot2) m <- microbenchmark::microbenchmark( "base 1u"=iris[order(iris$Petal.Length),], "sortdf 1u"=sortdf(iris,"Petal.Length"), "arrange 1u"=dplyr::arrange(iris,Petal.Length), "base 1d"=iris[order(iris$Petal.Length,decreasing=TRUE),], "sortdf 1d"=sortdf(iris,"Petal.Length",dir="down"), "arrange 1d"=dplyr::arrange(iris,-Petal.Length), "base 2d"=iris[order(iris$Petal.Length,iris$Sepal.Length,decreasing=TRUE),], "sortdf 2d"=sortdf(iris,c("Petal.Length","Sepal.Length"),dir=c("down","down")), "arrange 2d"=dplyr::arrange(iris,-Petal.Length,-Sepal.Length), "base 1u1d"=iris[order(iris$Petal.Length,rev(iris$Sepal.Length)),], "sortdf 1u1d"=sortdf(iris,c("Petal.Length","Sepal.Length"),dir=c("up","down")), "arrange 1u1d"=dplyr::arrange(iris,Petal.Length,-Sepal.Length), times=1000 ) autoplot(m)+theme_bw()
内容的提问来源于stack exchange,提问作者mindlessgreen
相关产品推荐
相关产品推荐

