使用case_when按类别值占比替换数据框国家名称的问题排查
问题排查与修正
原代码的核心错误
- 逻辑运算符误用:用
|(逻辑或)替代了&(逻辑与),导致只要是亚洲/欧洲国家就会被替换,完全忽略了占比条件。 - 阈值数值错误:需求是10%(
0.1),但代码中写的是0.01(1%),不符合要求。 - 缺少保留原国家的分支:没有设置占比达标时保留原
country的逻辑,所有行都被匹配到替换条件。
修正后的代码
library(dplyr) df_trial <- df %>% group_by(category) %>% mutate( # 计算当前国家value在该category中的占比 prop = value / sum(value, na.rm = TRUE), country = case_when( # 亚洲国家且占比<10%,替换为Other Asian countries continent == 'Asia' & prop < 0.1 ~ 'Other Asian countries', # 欧洲国家且占比<10%,替换为Other European countries continent == 'Europe' & prop < 0.1 ~ 'Other European countries', # 占比达标时保留原国家名称 TRUE ~ country ) ) %>% # 可选:如果不需要prop列,可以移除 select(-prop)
验证结果
运行后查看df_trial,会得到符合预期的结果:
- CHN(value=10)在每个category的占比约37.17%,≥10%,保留原名称。
- LAO(value=0.9)占比约3.35%,<10%,替换为
Other Asian countries。 - GBR(value=15)占比约55.76%,≥10%,保留原名称。
- DEU(value=1)占比约3.72%,<10%,替换为
Other European countries。
内容的提问来源于stack exchange,提问作者msug
相关产品推荐
相关产品推荐

