You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按组随机选取连续行的R语言技术实现问题

解决按学生随机选择连续考试记录的问题

问题原因

你写的代码里,1:counts$counts是长度为3的向量(对应3个学生的考试次数),R在处理sample(1:counts$counts, 1)时只会取向量的第一个元素(也就是5),导致所有学生的起始、结束值都是从1-5里随机选的同一个值,结果完全一致,同时触发警告。

正确解决方案

方法一:先生成每个学生的随机范围,再筛选数据

library(dplyr)

# 为每个学生生成随机的连续考试范围
student_ranges <- my_data %>%
  group_by(student_id) %>%
  summarise(total_exams = n()) %>%
  rowwise() %>% # 关键:逐行处理每个学生
  mutate(
    start = sample(1:total_exams, 1),
    end = sample(start:total_exams, 1)
  )

# 关联原数据并筛选符合范围的记录
filtered_data <- my_data %>%
  left_join(student_ranges, by = "student_id") %>%
  filter(exam_no >= start & exam_no <= end)

方法二:直接在分组内完成筛选(更简洁)

library(dplyr)

filtered_data <- my_data %>%
  group_by(student_id) %>%
  mutate(
    total_exams = n(),
    # 每个组内随机生成起始考试号
    start = sample(1:total_exams, 1),
    # 从起始号到最后一次考试里随机选结束号
    end = sample(start:total_exams, 1)
  ) %>%
  # 筛选出在范围内的考试记录
  filter(exam_no >= start & exam_no <= end) %>%
  # 移除临时生成的辅助列
  select(-total_exams, -start, -end)

验证效果

每次运行代码,每个学生的连续考试范围都会随机生成,示例结果如下(每次运行结果不同):

student_id exam_no   result
1           1       2 55.76655
2           1       3 53.81728
3           1       4 74.82202
4           2       1 58.32422
5           2       2 60.38213
6           3       3 37.97265
7           3       4 52.36436
8           3       5 61.62080

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 10:55:50