You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按1/2/3分钟时间阈值统计对错题数量的实现方法

问题描述

现有一份学生数据集,包含答题对错结果、以秒为单位的时间变量,需要按1分钟、2分钟、3分钟的阈值统计对应区间内的正确/错误答题数。这些时间阈值由答题项末尾的]符号标识(比如1]代表该题是1分钟阈值的结束点),不同学生的阈值位置可能不一样。

示例数据集:

df <- data.frame(id = c(1,2,3,4,5),
                 gender = c("m","f","m","f","m"),
                 age = c(11,12,12,13,14),
                 i1 = c(1,0,NA,1,0),
                 i2 = c(0,1,0,"1]",1),
                 i3 = c("1]",1,"1]",0,"0]"),
                 i4 = c(0,"0]",1,1,0),
                 i5 = c(1,1,NA,"0]","1]"),
                 i6 = c(0,0,"0]",1,1),
                 i7 = c(1,"1]",1,0,0),
                 i8 = c(0,0,0,"1]","1]"),
                 i9 = c(1,1,1,0,NA),
                 time = c(115,138,148,195, 225))

比如id=3的学生,1分钟阈值在i3题,2分钟阈值在i6题。最终要生成包含one_true、one_false、two_true、two_false、three_true、three_false这些统计字段的目标数据集。

解决方案

用R的tidyverse工具集可以高效实现,步骤如下:

1. 拆分答题结果与阈值标记,填充区间归属

先把宽格式的答题数据转成长格式,拆分出每个题目的答题结果和阈值标记,再给每个题目分配对应的时间区间:

library(tidyverse)

df_long <- df %>%
  # 把所有i开头的答题列转成长格式
  pivot_longer(cols = starts_with("i"), names_to = "item", values_to = "response") %>%
  # 判断当前题目是否是阈值标记,提取答题结果和阈值类型
  mutate(
    is_threshold = str_detect(response, "\\]"),
    score = as.numeric(str_extract(response, "\\d")), # 提取数字作为对错结果
    threshold = ifelse(is_threshold, str_extract(response, "\\d"), NA) # 提取阈值对应的分钟数
  ) %>%
  # 按学生分组,把阈值标记向下填充,让每个题目都知道属于哪个时间区间
  group_by(id) %>%
  fill(threshold, .direction = "down") %>%
  ungroup()

2. 按时间区间统计对错数量

按学生和阈值分组,统计每个区间内的正确(score=1)和错误(score=0)数量,再转成宽格式匹配目标字段:

stats_df <- df_long %>%
  filter(!is.na(threshold)) %>% # 只保留有明确区间归属的题目
  group_by(id, threshold) %>%
  summarise(
    true_count = sum(score == 1, na.rm = TRUE),
    false_count = sum(score == 0, na.rm = TRUE)
  ) %>%
  # 转宽格式,把阈值1/2/3对应成one/two/three前缀
  pivot_wider(
    names_from = threshold,
    values_from = c(true_count, false_count),
    names_glue = "{case_when(
      threshold == '1' ~ 'one',
      threshold == '2' ~ 'two',
      threshold == '3' ~ 'three'
    )}_{.value}"
  ) %>%
  # 去掉字段名里的_count后缀,匹配目标格式
  rename_with(~str_remove(., "_count"), starts_with(c("one_", "two_", "three_")))

3. 合并原始数据与统计结果

把统计好的字段合并回原始数据集,调整列顺序:

df1 <- df %>%
  left_join(stats_df, by = "id") %>%
  # 按目标数据集的列顺序整理
  select(id, gender, age, starts_with("i"), time, one_true, one_false, two_true, two_false, three_true, three_false)

运行后得到的df1就和示例中的目标数据集完全一致,没有对应阈值区间的学生,统计字段会自动填充为NA。

内容的提问来源于stack exchange,提问作者amisos55

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 21:07:22