You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在dplyr中按分组/筛选计算p值及简化统计流程?

Hey there! Let's break down your questions and find better ways to handle your data analysis. First, let's recap your current workflow: you're repeating a filter-group-summarise process 10 times for each type from 1 to 10. We can streamline this and add statistical tests easily.

Your Sample Data (for reference)

data <- read.table(header=T, text= ' PID time tdif testno type
3 205 0 1 1
4 77 0 1 1
4 85 8 2 1
4 126 41 3 1
4 165 39 4 1
4 202 37 5 1
4 238 36 6 1
4 272 34 7 1
4 277 5 8 1
4 370 93 9 1
4 397 27 10 1
4 452 55 11 1
4 522 70 12 1
4 529 7 13 1
4 608 79 14 1
4 651 43 15 1
4 655 4 16 1
4 713 58 17 1
4 804 91 18 1
4 900 96 19 1
4 944 44 20 1
4 979 35 21 1
4 1015 36 22 1
4 1051 36 23 1
4 1077 26 24 1
4 1124 47 25 1
4 1162 38 26 1
4 1222 60 27 1
4 1334 112 28 1
4 1383 49 29 1
4 1457 74 30 1
4 1506 49 31 1
4 1590 84 32 1
4 1768 178 33 1
4 1838 70 34 1
4 1880 42 35 1
4 1915 35 36 1
4 1973 58 37 1
4 2017 44 38 1
4 2090 73 39 1
4 2314 224 40 1
4 2381 67 41 1
4 2433 52 42 1
4 2484 51 43 1
4 2694 210 44 1
4 2731 37 45 1
4 2792 61 46 1
4 2958 166 47 1
5 48 0 1 3
5 111 63 2 3
5 699 588 3 3
5 1077 378 4 3
6 -43 0 1 3
8 67 0 1 1
8 168 101 2 1
8 314 146 3 1
8 368 54 4 1
8 586 218 5 1
10 639 0 1 6
13 -454 0 1 3
13 -384 70 2 3
13 -185 199 3 3
13 193 378 4 3
13 375 182 5 3
13 564 189 6 3
13 652 88 7 3
13 669 17 8 3
13 718 49 9 3
14 704 0 1 8
15 -165 0 1 3
15 -138 27 2 3
15 1335 1473 3 3
16 168 0 1 6
18 -1329 0 1 3
18 -1177 152 2 3
18 -1071 106 3 3
18 -945 126 4 3
18 -834 111 5 3
18 -719 115 6 3
18 -631 88 7 3
18 -497 134 8 3
18 -376 121 9 3
18 -193 183 10 3
18 -78 115 11 3
18 -13 65 12 3
18 100 113 13 3
18 196 96 14 3
18 552 356 15 3
18 650 98 16 3
18 737 87 17 3
18 804 67 18 3
18 902 98 19 3
18 983 81 20 3
18 1119 136 21 3
19 802 0 1 1
19 1593 791 2 1
26 314 0 1 8
26 389 75 2 8
26 597 208 3 8
33 639 0 1 6 ')

1. Simplify Your Summary Calculation

You don't need to repeat the process 10 times! Instead, group by both type and testno in one go. This will automatically calculate median, IQR, and sample size for every combination of type (1-10) and testno:

library(dplyr)

# One-step summary for all type 1-10
result_all <- data %>%
  filter(type %in% 1:10) %>% # Only include types 1 through 10
  group_by(type, testno) %>%
  summarise(
    Median = median(tdif, na.rm = TRUE), # Add na.rm to handle missing values (if any)
    IQR = IQR(tdif, na.rm = TRUE),
    n = n(),
    .groups = "drop" # Cleans up the grouping structure after summarising
  )

# View the result
head(result_all)

This gives you a tidy data frame with all your stats in one table, no repetitive code required.


2. Adding P-Value Tests for Groups

Since you're using median and IQR, it's likely your data isn't normally distributed—so non-parametric tests are a good fit. Below are common scenarios for testing differences between groups:

Scenario 1: Test differences across testno within each type

Use the Kruskal-Wallis test to check if there's a statistically significant difference in tdif across different testno values for each type. We can use the broom package to turn test results into a tidy data frame:

library(broom)

# Kruskal-Wallis test for each type
kruskal_by_type <- data %>%
  filter(type %in% 1:10) %>%
  group_by(type) %>%
  summarise(
    test_result = list(
      kruskal.test(tdif ~ factor(testno), data = cur_data()) %>%
        tidy() # Convert test output to a data frame
    ),
    .groups = "drop"
  ) %>%
  unnest(test_result)

# View the test results (includes p-value, chi-squared statistic, etc.)
kruskal_by_type

Scenario 2: Test differences across type within each testno

If you want to compare tdif across different type values for each testno, just swap the grouping variable:

# Kruskal-Wallis test for each testno
kruskal_by_testno <- data %>%
  filter(type %in% 1:10) %>%
  group_by(testno) %>%
  summarise(
    test_result = list(
      kruskal.test(tdif ~ factor(type), data = cur_data()) %>%
        tidy()
    ),
    .groups = "drop"
  ) %>%
  unnest(test_result)

Scenario 3: Pairwise comparisons (post-hoc tests)

If the Kruskal-Wallis test returns a significant p-value, you might want to do pairwise comparisons to see which groups differ. Use the dunn.test package for this (adjust for multiple comparisons with Bonferroni):

library(dunn.test)

# Pairwise Dunn tests for a single type (e.g., type 1)
pairwise_type1 <- data %>%
  filter(type == 1) %>%
  dunn.test(tdif, factor(testno), method = "bonferroni")

# To run this for all types at once:
pairwise_all_types <- data %>%
  filter(type %in% 1:10) %>%
  group_by(type) %>%
  summarise(
    pairwise_tests = list(
      dunn.test(tdif, factor(testno), method = "bonferroni") %>%
        as.data.frame() # Convert results to a data frame for tidying
    ),
    .groups = "drop"
  )

内容的提问来源于stack exchange,提问作者JannickLinden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:09:46