You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr计算基于前两次考试结果的当前考试条件概率

解决方案:高效计算考试结果的条件概率

核心思路

条件概率P(当前考试=k | 前一次=a, 前两次=b)的计算逻辑为:该(前两次=a,前一次=b,当前=k)组合的出现频次 ÷ (前两次=a,前一次=b)组合的总频次。通过dplyr的分组统计+格式转换可快速生成所有8种组合的标准化结果。

代码实现(基于dplyr)

假设你已通过data.table::shift生成包含current_exam(当前考试结果)、prev1_exam(前一次考试结果)、prev2_exam(前两次考试结果)的数据集exam_sequences,执行以下步骤:

library(dplyr)
library(tidyr)

# 过滤无有效前两次记录的行(排除NA值)
filtered_data <- exam_sequences %>%
  filter(!is.na(prev1_exam), !is.na(prev2_exam))

# 计算条件概率并整理为8行标准化格式
conditional_probs <- filtered_data %>%
  # 按前两次考试结果分组(条件部分)
  group_by(prev2_exam, prev1_exam) %>%
  # 统计每组总次数及当前考试各结果的频次
  summarise(
    total_count = n(),
    count_0 = sum(current_exam == 0),
    count_1 = sum(current_exam == 1),
    .groups = "drop"
  ) %>%
  # 计算条件概率
  mutate(
    prob_0 = count_0 / total_count,
    prob_1 = count_1 / total_count
  ) %>%
  # 转换为窄表格式,每种组合对应一行
  pivot_longer(
    cols = starts_with("prob_"),
    names_to = "current_exam",
    values_to = "probability",
    names_prefix = "prob_"
  ) %>%
  # 修正当前考试结果的数值类型
  mutate(current_exam = as.integer(current_exam)) %>%
  # 排序以提升可读性
  arrange(prev2_exam, prev1_exam, current_exam)

代码正确性验证

  1. 概率和校验:每个(prev2_exam, prev1_exam)分组对应的两个概率之和应为1(允许微小浮点误差),可通过以下代码检查:
    conditional_probs %>%
      group_by(prev2_exam, prev1_exam) %>%
      summarise(total_prob = sum(probability), .groups = "drop")
    
  2. 手动对比校验:任选一组(比如prev2=0, prev1=0),手动统计该组合的总次数及current_exam=0的次数,计算概率后与代码输出对比,验证一致性。

标准化结果示例

最终输出的8行结果格式如下:

prev2_examprev1_examcurrent_examprobability
0000.85
0010.15
0100.60
0110.40
1000.70
1010.30
1100.45
1110.55

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:40:54