titre水平趋势分类:放宽1个偏差判定规则的R语言实现求助
实现允许单个偏差值的趋势分类
需求概述
- 原趋势分类规则:
- 数值单调递增(含持平)→
uptrend - 数值单调递减(含持平)→
downtrend - 所有数值完全相同→
stable - 其余情况→
fluctuate
- 数值单调递增(含持平)→
- 放宽规则:允许存在1个偏差值,即移除最多1个元素后,剩余数据符合原分类逻辑的,按对应类别标记。例如:
- id=9的序列
1,2,1,1,1,1,1,1,1移除唯一的2后全为1,归为stable - id=10-12的样本同理,按移除1个偏差值后的结果分类
- id=9的序列
修改后的代码
library(tidyverse) # 定义判断序列是否符合单调递增(含持平)的函数 is_uptrend <- function(x) all(diff(x) >= 0) # 定义判断序列是否符合单调递减(含持平)的函数 is_downtrend <- function(x) all(diff(x) <= 0) # 定义判断序列是否全相同的函数 is_stable <- function(x) all(x == x[1]) # 处理数据 df_processed <- df %>% # 将titre_level转换为数值向量列表 mutate(titre_numeric = map(str_split(titre_level, ","), as.numeric)) %>% group_by(id) %>% mutate( # 先判断原序列是否符合原规则 original_stable = is_stable(titre_numeric[[1]]), original_uptrend = is_uptrend(titre_numeric[[1]]), original_downtrend = is_downtrend(titre_numeric[[1]]), # 检查移除1个元素后是否能变为stable:统计每个值的出现次数,最多的那个的数量 >= 总长度-1 can_be_stable = max(table(titre_numeric[[1]])) >= length(titre_numeric[[1]]) - 1, # 检查移除1个元素后是否能变为uptrend:遍历每个位置,移除后检查是否符合uptrend can_be_uptrend = any(map_lgl(seq_along(titre_numeric[[1]]), ~is_uptrend(titre_numeric[[1]][-.x]))), # 检查移除1个元素后是否能变为downtrend:遍历每个位置,移除后检查是否符合downtrend can_be_downtrend = any(map_lgl(seq_along(titre_numeric[[1]]), ~is_downtrend(titre_numeric[[1]][-.x]))) ) %>% ungroup() %>% # 确定最终趋势分类 mutate( trend2 = case_when( original_stable | can_be_stable ~ "stable", original_uptrend | can_be_uptrend ~ "uptrend", original_downtrend | can_be_downtrend ~ "downtrend", TRUE ~ "fluctuate" ) ) %>% # 选择需要的列 select(id, titre_level, trend, trend2) # 查看结果 print(df_processed)
代码解释
- 辅助函数定义:提前定义
is_uptrend、is_downtrend、is_stable三个函数,分别用于判断序列是否符合单调递增、单调递减、全相同的规则,简化后续逻辑。 - 数据转换:将每个id的
titre_level字符串拆分为数值向量列表,方便后续处理。 - 多维度检查:
- 先判断原序列是否符合原规则;
can_be_stable:通过统计值的出现频率,判断是否存在某个值的数量占比≥总长度-1(即最多只有1个偏差值);can_be_uptrend/can_be_downtrend:遍历移除序列中的每个元素,检查是否存在移除后符合单调递增/递减的情况。
- 最终分类:按优先级判断:优先标记
stable(原规则或可通过移除1个值变为stable),其次是uptrend,再是downtrend,最后是fluctuate。
测试数据
df <- structure(list(id = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12), titre_level = c("2,1,1,1,1", "4,4,4,4,4,1,1,1,1,1", "8,6,6,6,6,4", "1,1,1,1,1,1,1,1,1,3,3", "4,4,7,10", "8,11", "1,1", "6,6,6,6,6,6,6,6,6", "1,2,1,1,1,1,1,1,1", "8,8,6,8", "4,2,2,2,4,2,2,2,1", "1,1,2,4,4,1" ), trend = c("downtrend", "downtrend", "downtrend", "uptrend", "uptrend", "uptrend", "stable", "stable", "stable", "stable", "downtrend", "uptrend")), class = "data.frame", row.names = c(NA, 12L))
内容的提问来源于stack exchange,提问作者HNSKD
相关产品推荐
相关产品推荐

