dplyr中group_by+summarize结合case_when生成重复行问题的解决咨询
case_when() return duplicate rows in summarize() while ifelse() doesn't? Hey there, let's break down why you're seeing duplicate rows with case_when() but not with ifelse(), and how to fix it to get the unique-row result you want directly in summarize().
The Issue Recap
When using group_by(id) %>% summarize() with case_when(), groups with multiple rows end up repeating the 'Multiple Colors' value once per original row in the group. But ifelse() gives you one row per group as expected. Even weirder, when both case_when() branches return hardcoded strings, it works correctly too.
What's Going On?
The root cause is how case_when() handles vector lengths compared to ifelse():
ifelse()behavior: It returns a vector that matches the length of the condition argument. When you usen() > 1as the condition, that's a length-1 logical value (one per group). Even if thefav_colorargument is longer (e.g., 2 elements for id=1),ifelse()just returns a length-1 result ('Multiple Colors') for the group, whichsummarize()uses to create one row per group.case_when()behavior: It aligns all inputs to the longest vector in any of its branches. For groups with multiple rows,fav_coloris a vector of lengthn()(e.g., 2 for id=1). When you writen() > 1 ~ 'Multiple Colors',case_when()repeats that length-1 string to match the length offav_color(2 times). Sincesummarize()expands any length>1 result into multiple rows, you get one row per original row in the group.
When both branches return hardcoded strings, they're both length-1, so case_when() returns a length-1 vector per group—hence no duplicates.
Fixes to Get Unique Rows with case_when()
1. Force Length-1 Output in the n() == 1 Branch
Instead of returning the full fav_color vector, explicitly return a single value using first(), last(), or indexing with [1]:
dat %>% group_by(id) %>% summarize( fav_color = case_when( n() > 1 ~ 'Multiple Colors', n() == 1 ~ first(fav_color) ) )
This ensures both branches return a length-1 vector, so summarize() creates exactly one row per group.
2. Stick with ifelse() (Your Working Example)
Since you already found that ifelse() behaves as expected, this is a totally valid solution. It's less flexible than case_when() for complex logic, but it works perfectly here.
3. Combine mutate() + slice_head()
If you prefer to use case_when() without modifying branch outputs, first mutate the column for all rows in the group, then keep just one row per group:
dat %>% group_by(id) %>% mutate( fav_color = case_when( n() > 1 ~ 'Multiple Colors', TRUE ~ fav_color ) ) %>% slice_head(n = 1) %>% ungroup()
Expected Result
All these approaches will give you your desired unique-row output:
# A tibble: 5 x 2 id fav_color <dbl> <chr> 1 1 Multiple Colors 2 2 red 3 3 green 4 4 Multiple Colors 5 5 pink
内容的提问来源于stack exchange,提问作者Tyler Rinker

