基于dplyr分组的条件匹配:返回对应ProjectID的技术问询
Hey there! I’ve tackled a similar grouping and matching task with dplyr before, so let me break down how to solve this for you. The key is to work within each cluster group to check for the presence of both projects A and C, then map C’s ProjectID to A’s Dependent column.
Step 1: Start with Example Data (for testing)
First, let’s create a sample dataset that mirrors your scenario—this makes it easier to test the code:
library(dplyr) df <- tibble( cluster = c("Cluster1", "Cluster1", "Cluster2", "Cluster2", "Cluster3"), Project = c("A", "C", "A", "B", "C"), ProjectID = c("A001", "C001", "A002", "B001", "C002"), Dependent = NA_character_ )
Step 2: The dplyr Workflow
Here’s the code to implement your logic. I’ll explain each part so you understand what’s happening:
df_processed <- df %>% # Group the data by cluster so we operate within each cluster's context group_by(cluster) %>% mutate( # First, check if the current cluster has both Project A and Project C has_both_projects = all(c("A", "C") %in% Project), # Extract ProjectID of Project C (only if both projects exist in the cluster) c_id = ifelse(has_both_projects, ProjectID[Project == "C"], NA), # Assign C's ID to Dependent column ONLY for rows where Project is A AND both projects exist Dependent = ifelse(Project == "A" & has_both_projects, c_id, Dependent) ) %>% # Clean up: remove the temporary helper columns (optional but keeps data tidy) select(-has_both_projects, -c_id) %>% # Ungroup to return to a regular tibble (always good practice after grouping) ungroup()
Step 3: Check the Result
If you run print(df_processed), you’ll see the desired output:
#> # A tibble: 5 × 4 #> cluster Project ProjectID Dependent #> <chr> <chr> <chr> <chr> #> 1 Cluster1 A A001 C001 #> 2 Cluster1 C C001 NA #> 3 Cluster2 A A002 NA #> 4 Cluster2 B B001 NA #> 5 Cluster3 C C002 NA
Edge Case Handling
What if a cluster has multiple instances of Project C? The code above will grab the first C’s ProjectID. If you want to combine all C IDs (e.g., separated by commas), modify the c_id line like this:
c_id = ifelse(has_both_projects, paste0(ProjectID[Project == "C"], collapse = ", "), NA)
That’s it! This approach keeps everything within the dplyr pipeline and clearly implements your grouping and matching logic.
内容的提问来源于stack exchange,提问作者OBoro

