R语言:解决DataFrame中循环计算迭代次数新增列的报错问题
计算DataFrame中迭代调整次数的正确实现方法
问题背景
给定如下DataFrame:
dataf <- data.frame(category = c("Above", "Below", "In Range", "Above"), A = c(1.99, 4.99, 6.99, 5.99), B = c(0.99, 9.99, 6.99, 1.99))
需要新增一列C,规则如下:
- 当
category为"Above":每次将A减少20%,直到A小于B×1.1,记录迭代次数 - 当
category为"Below":每次将A增加20%,直到A≥B,记录迭代次数 - 当
category为"In Range":记录0次
期望输出:
desired_dataf <- data.frame(category = c("Above", "Below", "In Range", "Above"), A = c(1.99, 4.99, 6.99, 5.99), B = c(0.99, 9.99, 6.99, 1.99), C = c(3, 4, 0, 5))
错误原因分析
你编写的自定义函数是标量函数,仅能处理单个数值,但在mutate中传入的是整个列(向量),while循环无法处理长度大于1的条件判断,因此报错:the condition has length > 1。单独测试时传入单个值能正常运行,但批量处理向量时就会出错。
解决方案
方法1:将标量函数向量化
用Vectorize()包装自定义函数,让它支持向量输入:
library(dplyr) # 原自定义函数不变 decrease_func <- function(x, y, count=0, max=1.1) { new_x <- x while((new_x/y) > max) { new_x = new_x * (.8) count = count + 1 } return(count) } increase_func <- function(x, y, count=0, min=1.0) { new_x <- x while((new_x / y) < min) { new_x = new_x * (1.2) count = count + 1 } return(count) } # 向量化函数 vec_decrease <- Vectorize(decrease_func) vec_increase <- Vectorize(increase_func) # 生成结果 new_df <- dataf %>% mutate(C = case_when( category == "Above" ~ vec_decrease(A, B), category == "Below" ~ vec_increase(A, B), TRUE ~ 0L )) print(new_df)
方法2:用rowwise()逐行处理
通过rowwise()让mutate逐行执行函数,保持原函数的标量形式:
library(dplyr) # 原自定义函数不变 decrease_func <- function(x, y, count=0, max=1.1) { new_x <- x while((new_x/y) > max) { new_x = new_x * (.8) count = count + 1 } return(count) } increase_func <- function(x, y, count=0, min=1.0) { new_x <- x while((new_x / y) < min) { new_x = new_x * (1.2) count = count + 1 } return(count) } # 逐行处理 new_df <- dataf %>% rowwise() %>% mutate(C = case_when( category == "Above" ~ decrease_func(A, B), category == "Below" ~ increase_func(A, B), TRUE ~ 0L )) %>% ungroup() # 取消逐行模式 print(new_df)
方法3:用数学公式直接计算(高效无循环)
迭代过程是等比数列,可以通过对数计算直接得到次数,避免循环,大数据量下效率更高:
library(dplyr) new_df <- dataf %>% mutate(C = case_when( category == "Above" ~ { ratio <- A / (1.1 * B) # 如果初始已经满足条件,返回0;否则计算需要的次数,向上取整 ifelse(ratio <= 1, 0, ceiling(log(1/ratio) / log(1/0.8))) }, category == "Below" ~ { ratio <- B / A # 如果初始已经满足条件,返回0;否则计算需要的次数,向上取整 ifelse(ratio <= 1, 0, ceiling(log(ratio) / log(1.2))) }, TRUE ~ 0L )) print(new_df)
验证:该方法计算结果与期望输出完全一致,且无需循环,运算速度更快。
内容的提问来源于stack exchange,提问作者Brian
相关产品推荐
相关产品推荐

