You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按分类控制真假比例新增随机布尔值列的实现方法

基础R实现方案

利用ave()函数按category分组,每组内生成等量的TRUE/FALSE后打乱顺序:

df$random_bool <- ave(df$category, df$category, FUN = function(group) {
  # 每组生成占比1:1的布尔值后随机打乱
  sample(rep(c(TRUE, FALSE), each = length(group) / 2))
})

如果分组行数为奇数,可调整为如下写法适配:

sample(c(rep(TRUE, ceiling(length(group)/2)), rep(FALSE, floor(length(group)/2))))

dplyr/tidyverse实现方案

分组后链式操作,可读性更高:

library(dplyr)

df <- df %>%
  group_by(category) %>%
  mutate(random_bool = sample(rep(c(TRUE, FALSE), each = n() / 2))) %>%
  ungroup()

结果验证

执行如下代码即可确认每个分类下的TRUE/FALSE数量符合要求:

table(df$category, df$random_bool)

输出示例:

FALSE TRUE
  NEGATIVE      5    5
  NEUTRAL       5    5
  POSITIVE      5    5

内容的提问来源于stack exchange,提问作者Madamadam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 15:24:03