如何在R中使用正则提取特殊字符前后内容并创建数据框
Got it, let's tackle this problem step by step! You've already nailed extracting numerical values, but the missing piece is linking those numbers to their corresponding regions and standardizing everything into a tidy data frame. Here's how to make that happen:
Step 1: Update Your Regex to Capture Both Regions and Values
Instead of just grabbing numbers, we need a regex that matches the full [Region] = [Value] pattern. This lets us pair each number with its associated region label. The regex below handles variations like spaces around the equals sign, hyphenated regions, and abbreviations:
pattern <- "([A-Za-z\\.\\-\\s]+)\\s*=\\s*([0-9]+\\.*[0-9]*)"
([A-Za-z\\.\\-\\s]+): Captures the region name (covers letters, dots, hyphens, and spaces)\\s*=\\s*: Matches the equals sign with optional spaces on either side([0-9]+\\.*[0-9]*): Captures the numeric value (works for integers and decimals)
Step 2: Build a Helper Function to Parse Individual Strings
We'll create a function that takes one string from your vector, extracts all region-value pairs, cleans up the region names, and returns a named list (which will become a row in our data frame):
library(stringr) library(purrr) parse_string <- function(str) { # Extract all region-value matches from the string matches <- str_match_all(str, pattern)[[1]] # Clean and standardize region names regions <- matches[, 2] %>% str_trim() %>% # Remove leading/trailing spaces str_to_title() %>% # Capitalize first letter of each word str_replace_all(c("\\s+" = "_", "\\-" = "_", "\\." = "_")) # Replace spaces/hyphens/dots with underscores # Convert extracted values to numeric type values <- as.numeric(matches[, 3]) # Return as a named list (ready to become a data frame row) set_names(as.list(values), regions) }
Step 3: Process the Entire Vector and Build the Data Frame
Use map_dfr from the purrr package to apply our function to every element in VEC and automatically bind the results into a data frame. Any regions missing from a row will be filled with NA:
# Test with your sample data VEC <- c( "today's performance: SW= 33.9; west =13.85 East=70.09 and the north central =125.823, S.S. = 41.5", "sent as received South= 28.76 while the West =44.55, SouthEast=64.32; north-east =130; south west = 38.86" ) # Generate the structured data frame df <- map_dfr(VEC, parse_string) # View the final result print(df)
Sample Output:
Sw West East North_Central S_S South SouthEast North_East South_West 1 33.9 13.85 70.09 125.823 41.5 NA NA NA NA 2 NA 44.55 NA NA NA 28.76 64.32 130 38.86
Optional: Standardize Region Names Even More
If you want to map abbreviations (like Sw to South_West, S_S to South_South) to consistent, human-readable names, add a lookup step in the parse_string function:
# Define a lookup for consistent region names region_lookup <- c( "Sw" = "South_West", "S_S" = "South_South", "North_Central" = "North_Central", "North_East" = "North_East", "SouthEast" = "South_East", "SouthWest" = "South_West" ) # Update the region cleaning step in parse_string regions <- matches[, 2] %>% str_trim() %>% str_to_title() %>% str_replace_all(c("\\s+" = "_", "\\-" = "_", "\\." = "_")) %>% recode(!!!region_lookup)
This ensures your column names are perfectly consistent, even if the input uses different abbreviations or spellings.
内容的提问来源于stack exchange,提问作者Tunde Awosanya

