如何使用dplyr包的recode函数重新编码含括号的字符串值?
Hey AleksP, let's break down what's going on with your recode issue and fix it with more robust, maintainable approaches!
First, the core reason some of your recode rules aren't working is that dplyr::recode relies on exact, case-sensitive matches. Even tiny, easy-to-miss differences—like hidden leading/trailing spaces, non-breaking spaces, or inconsistent punctuation—will make a rule fail. For example, if your raw data has "RKS (UNMIK)" (no leading space) but your rule uses " RKS (UNMIK)" (with a leading space), it won't match. That's almost certainly why "EGY. EU" isn't being replaced—double-check the raw value is exactly the string you're using in your code (run unique(data$parties) to list all unique values and verify).
As for the weird bracket "matching" between "(CHE" and "NOR)", that's not recode doing any automatic bracket pairing—it's likely a coincidence in your data, or leftover values from partial edits elsewhere. Let's fix this properly with better methods.
Better Approaches for Recoding ISO3 Country Codes
1. Mapping Table + left_join (Most Maintainable)
This is my go-to for recoding with multiple rules—it keeps your mappings organized, easy to update, and avoids exact-match typos.
First, create a clear table of your raw-to-target mappings:
library(dplyr) # Define your raw values and their target ISO3 codes code_mappings <- tibble( raw_parties = c(" RKS (UNMIK)", "(CHE", "EGY. EU", "BRA. PRY", "CU", "NOR)", "VNM)KOR"), target_iso3 = c("RKS", "CHE", "EGY,EU", "BRA-PRY", "CUB", "NOR", "VNM,KOR") ) # Join the mapping to your data and replace values data <- data %>% left_join(code_mappings, by = c("parties" = "raw_parties")) %>% # Use coalesce to keep original values when no match exists mutate(parties = coalesce(target_iso3, parties)) %>% # Clean up the temporary mapping column select(-target_iso3)
2. case_when + Regular Expressions (For Flexible Matching)
If you need to handle variations (like extra spaces or inconsistent punctuation), case_when paired with stringr functions lets you use regex to match patterns instead of rigid exact strings:
library(dplyr) library(stringr) data <- data %>% mutate(parties = case_when( # Match any variation of RKS (UNMIK) with optional leading spaces str_detect(parties, "^\\s*RKS \\(UNMIK\\)$") ~ "RKS", # Remove leading bracket from "(CHE" str_detect(parties, "^\\(CHE$") ~ "CHE", # Remove trailing bracket from "NOR)" str_detect(parties, "^NOR\\)$") ~ "NOR", # Replace dot with comma in "EGY. EU" str_detect(parties, "^EGY\\. EU$") ~ "EGY,EU", # Replace dot with dash in "BRA. PRY" str_detect(parties, "^BRA\\. PRY$") ~ "BRA-PRY", # Exact match for "CU" parties == "CU" ~ "CUB", # Fix the mixed bracket in "VNM)KOR" str_detect(parties, "^VNM\\)KOR$") ~ "VNM,KOR", # Keep all other values as-is TRUE ~ parties ))
3. Pre-Clean Strings First
Before recoding, fix common messy data issues to reduce the number of rules you need:
data <- data %>% mutate( # Trim leading/trailing spaces from all entries parties = str_trim(parties), # Remove any leading or trailing brackets (adjust regex if needed) parties = str_remove_all(parties, "^\\(|\\)$") )
Quick Troubleshooting Tip
To confirm exactly what's in your parties column (and spot hidden characters), run:
# List all unique values to cross-check with your rules unique(data$parties) # View every character in each string to catch hidden spaces/punctuation stringr::str_view_all(data$parties, ".")
This will help you catch tiny discrepancies between your raw data and recode rules that you might have overlooked.
内容的提问来源于stack exchange,提问作者AleksP

