如何用dplyr按ICD_Grouping分组统计Immunohistochemistry各水平频数
Hey there! Let's get your dplyr code sorted to produce that wide-format data frame you need. Here's a step-by-step breakdown with working code:
Step 1: Fix the Grouping & Counting
Your initial code only grouped by ICD_Grouping, but we need to group by both ICD_Grouping and Immunohistochemistry to capture counts for each factor level (Yes/No/N/A) within each ICD group. We can use dplyr::count() as a handy shortcut for group_by() + summarise(n()):
Step 2: Reshape to Wide Format
Instead of struggling with spread() (which is now soft-deprecated in the tidyverse), we'll use tidyr::pivot_wider()—it's more intuitive and flexible. If you still want to use spread(), I'll include that version too.
Recommended: Using pivot_wider()
library(dplyr) library(tidyr) null.ta <- dbdata %>% filter(MutGroup == "Null") %>% # Count occurrences of each Immunohistochemistry level per ICD_Grouping count(ICD_Grouping, Immunohistochemistry) %>% # Reshape to wide format, filling missing levels with 0 pivot_wider( names_from = Immunohistochemistry, # Which column becomes new column names values_from = n, # Which column provides the values values_fill = 0 # Fill empty cells with 0 )
Legacy: Using spread()
If you prefer to stick with spread():
null.ta <- dbdata %>% filter(MutGroup == "Null") %>% group_by(ICD_Grouping, Immunohistochemistry) %>% summarise(count = n(), .groups = "drop") %>% # Drop grouping after summarizing spread( key = Immunohistochemistry, value = count, fill = 0 )
Step 3: Clean Up Unwanted Columns
If your data has extra factor levels (like that 4th level you mentioned), just use select() to keep only the columns you need. Note that N/A has special characters, so we wrap it in backticks:
null.ta <- null.ta %>% select(ICD_Grouping, Yes, No, `N/A`)
Final Output
This will give you exactly the format you're looking for:
| ICD_Grouping | Yes | No | N/A |
|---|---|---|---|
| C22 | 2 | 1 | 0 |
| C45 | 7 | 3 | 1 |
| C69 | 4 | 0 | 0 |
内容的提问来源于stack exchange,提问作者Seb Walpole

