R语言中非标准化州名向量转换为州缩写的实现问题
Got it, let's fix this—your vector of partial/shortened state names can absolutely be mapped to state.abb with some targeted string matching. The key is handling case insensitivity and partial prefix matches, since your entries are either full lowercase names or common shorthands. Here's a practical, robust solution:
Step 1: Prep for Case-Insensitive Matching
First, convert official state names to lowercase to avoid capitalization conflicts:
state_names_lower <- tolower(state.name)
Step 2: Build a Mapping Function
We'll create a function that checks each messy entry against official state names, finds valid prefix matches, and returns the corresponding abbreviation. We'll also add handling for ambiguous cases (like "miss" matching both Mississippi and Missouri) with manual overrides.
Option 1: Using stringr (cleaner syntax)
If you have the stringr package installed:
library(stringr) map_to_abb <- function(entry) { # Manual override for ambiguous shorthands (adjust as needed) manual_maps <- c("miss" = "MS") if (entry %in% names(manual_maps)) { return(manual_maps[entry]) } # Find indices of state names starting with the entry match_indices <- which(str_detect(state_names_lower, paste0("^", entry))) # Return abbreviation if exactly one match exists, else NA if (length(match_indices) == 1) { return(state.abb[match_indices]) } else { return(NA_character_) } }
Option 2: Base R (no extra packages)
If you prefer sticking to base R, use grep instead of stringr:
map_to_abb_base <- function(entry) { # Manual override for ambiguous shorthands manual_maps <- c("miss" = "MS") if (entry %in% names(manual_maps)) { return(manual_maps[entry]) } match_indices <- grep(paste0("^", entry), state_names_lower) if (length(match_indices) == 1) { return(state.abb[match_indices]) } else { return(NA_character_) } }
Step 3: Apply the Function to Your Vector
Run the function on your bs vector:
bs <- c("texas", "tex", "calif", "wisc", "mass", "miss", "oh", "ohio", "colo", "fla") # Using stringr version sapply(bs, map_to_abb, USE.NAMES = FALSE) # Or base R version sapply(bs, map_to_abb_base, USE.NAMES = FALSE)
Expected Output
Both methods will return the correct abbreviations:
[1] "TX" "TX" "CA" "WI" "MA" "MS" "OH" "OH" "CO" "FL"
Edge Case Notes
- Ambiguous Entries: The manual override handles cases where a shorthand matches multiple states (like "miss" for Mississippi/Missouri). You can expand this list with other ambiguous shorthands as needed.
- No Matches: The function returns
NAfor entries that don't match any state name prefix—you can adjust this to return a custom value (like "Unknown") if preferred.
内容的提问来源于stack exchange,提问作者jvalenti

