基于条件替换分组值:R语言dplyr处理水果数据问题求助
Replace "X" with Alternating Fruits in Grouped dplyr Data Frame
Got it, let's work through this problem together—you’re already off to a great start with grouping by Person using dplyr! Here’s how to wrap up the task of replacing those 'X' values with alternating fruits (so each person has one Apple and one Orange across days):
First, let's recap your starting data for clarity:
library(dplyr) # Original dataset df <- data.frame( Person=c("Ada","Ada","Bob","Bob","Carl","Carl"), Day=c(1,2,2,1,1,2), Fruit=c("Apple","X","Apple","X","X","Orange") )
Core Approach
The key idea is straightforward:
- For each person, identify their known fruit (the non-"X" value in their group)
- Replace "X" with the opposite fruit (Apple ↔ Orange) to maintain the alternation rule
Step-by-Step Implementation
Here’s a clean, readable way to do this with dplyr:
df_fixed <- df %>% group_by(Person) %>% # Create a helper column to store the person's known, non-X fruit mutate(known_fruit = Fruit[Fruit != "X"][1]) %>% # Replace X with the opposite fruit of the known value mutate(Fruit = case_when( Fruit != "X" ~ Fruit, # Keep existing non-X values as-is known_fruit == "Apple" ~ "Orange", known_fruit == "Orange" ~ "Apple" )) %>% # Remove the helper column (optional, if you don't need it) select(-known_fruit) %>% ungroup() # View the final result df_fixed
What This Does
group_by(Person): Ensures all operations stay isolated to each individual's dataknown_fruit = Fruit[Fruit != "X"][1]: Grabs the first (and only, in your dataset) non-"X" fruit for each personcase_when: Handles the replacement logic cleanly without messy nested conditionalsungroup(): Resets the data frame to its ungrouped state (good practice after grouping operations)
Handling Out-of-Order Days
If you want to explicitly ensure the alternation matches Day 1 and Day 2 (regardless of how the rows are ordered in the original data), add an arrange(Day) step inside the group:
df_fixed_ordered <- df %>% group_by(Person) %>% arrange(Day) %>% # Sort rows by Day first to lock in order mutate(known_fruit = Fruit[Fruit != "X"][1]) %>% mutate(Fruit = case_when( Fruit != "X" ~ Fruit, known_fruit == "Apple" ~ "Orange", known_fruit == "Orange" ~ "Apple" )) %>% select(-known_fruit) %>% ungroup()
Final Result
Running either code will give you the desired output:
| Person | Day | Fruit |
|---|---|---|
| Ada | 1 | Apple |
| Ada | 2 | Orange |
| Bob | 2 | Apple |
| Bob | 1 | Orange |
| Carl | 1 | Apple |
| Carl | 2 | Orange |
内容的提问来源于stack exchange,提问作者user186720
相关产品推荐
相关产品推荐

