R语言实现DataFrame逐行运行自定义函数并新增di_Flex列的技术问询
R语言for循环实现方案
前置说明
- 结合你提供的
sum_differences函数的运算逻辑,入参a实际对应同Company ID分组下的Flexibility_Thinking列,如果确需传入Company ID作为入参a,可自行修改下方代码中a的取值逻辑即可 - 实现逻辑为按唯一的
Company ID遍历循环计算,同组所有行共享同一个计算结果,无需逐行重复计算
完整可运行代码
# 加载依赖包 library(tidyverse) # 样例数据 data <- df <- data.frame(structure(list(`Company ID` = c(0, 0, 0, 0, 0, 5, 5, 5, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6), Action = c(5, 5, 2, 4, 5, 5, 3, 1, 5, 7, 2, 4, 2, 6, 2, 3, 1, 4, 1, 5), Flexibility_Thinking = c(7, 2, 5, 1, 6, 5, 7, 7, 4, 7, 5, 2, 3, 3, 3, 3, 4, 3, 6, 6), max_di_Flex = c(12.8, 12.8, 12.8, 12.8, 12.8, 8, 8, 8, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16)))) # 数据预处理 data <- data %>% group_by(`Company ID`) %>% filter(length(`Company ID`) > 1) data <- data %>% drop_na(`Flexibility_Thinking`) # 自定义函数 sum_differences <- function(a,b) { a <- unique(a) new_list <- c() for (i in a) { for (j in a) { if(i != j) { new_list <- c(new_list, abs(i-j)) } } } outcome <- round((sum(new_list) / length(a)), 2) percent <- outcome/b return(percent) } # 初始化新列 data$di_Flex <- NA # 获取所有唯一的公司ID unique_cid <- unique(data$`Company ID`) # for循环实现分组计算赋值 for (cid in unique_cid) { # 筛选当前公司对应的行索引 target_rows <- which(data$`Company ID` == cid) # 取当前分组的Flexibility_Thinking作为入参a param_a <- data$Flexibility_Thinking[target_rows] # 取当前分组的max_di_Flex(同组值一致,取第一个即可)作为入参b param_b <- data$max_di_Flex[target_rows][1] # 调用自定义函数计算结果 calc_res <- sum_differences(param_a, param_b) # 给当前分组所有行的di_Flex赋值 data$di_Flex[target_rows] <- calc_res } # 查看结果 head(data)
运行结果说明
运行完成后data数据集会新增di_Flex列,同Company ID的行结果一致,样例数据运算后:
- Company ID为0的行di_Flex值为1
- Company ID为5的行di_Flex值为0.25
- Company ID为6的行di_Flex值约为0.73
内容的提问来源于stack exchange,提问作者Theresa_S
相关产品推荐
相关产品推荐

