You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中使用正则提取特殊字符前后内容并创建数据框

Extract Structured Data Frame from String Vector with Region-Value Pairs

Got it, let's tackle this problem step by step! You've already nailed extracting numerical values, but the missing piece is linking those numbers to their corresponding regions and standardizing everything into a tidy data frame. Here's how to make that happen:

Step 1: Update Your Regex to Capture Both Regions and Values

Instead of just grabbing numbers, we need a regex that matches the full [Region] = [Value] pattern. This lets us pair each number with its associated region label. The regex below handles variations like spaces around the equals sign, hyphenated regions, and abbreviations:

pattern <- "([A-Za-z\\.\\-\\s]+)\\s*=\\s*([0-9]+\\.*[0-9]*)"
  • ([A-Za-z\\.\\-\\s]+): Captures the region name (covers letters, dots, hyphens, and spaces)
  • \\s*=\\s*: Matches the equals sign with optional spaces on either side
  • ([0-9]+\\.*[0-9]*): Captures the numeric value (works for integers and decimals)

Step 2: Build a Helper Function to Parse Individual Strings

We'll create a function that takes one string from your vector, extracts all region-value pairs, cleans up the region names, and returns a named list (which will become a row in our data frame):

library(stringr)
library(purrr)

parse_string <- function(str) {
  # Extract all region-value matches from the string
  matches <- str_match_all(str, pattern)[[1]]
  
  # Clean and standardize region names
  regions <- matches[, 2] %>%
    str_trim() %>%  # Remove leading/trailing spaces
    str_to_title() %>%  # Capitalize first letter of each word
    str_replace_all(c("\\s+" = "_", "\\-" = "_", "\\." = "_"))  # Replace spaces/hyphens/dots with underscores
  
  # Convert extracted values to numeric type
  values <- as.numeric(matches[, 3])
  
  # Return as a named list (ready to become a data frame row)
  set_names(as.list(values), regions)
}

Step 3: Process the Entire Vector and Build the Data Frame

Use map_dfr from the purrr package to apply our function to every element in VEC and automatically bind the results into a data frame. Any regions missing from a row will be filled with NA:

# Test with your sample data
VEC <- c(
  "today's performance: SW= 33.9; west =13.85 East=70.09 and the north central =125.823, S.S. = 41.5",
  "sent as received South= 28.76 while the West =44.55, SouthEast=64.32; north-east =130; south west = 38.86"
)

# Generate the structured data frame
df <- map_dfr(VEC, parse_string)

# View the final result
print(df)

Sample Output:

Sw  West  East North_Central S_S South SouthEast North_East South_West
1 33.9 13.85 70.09       125.823 41.5    NA        NA         NA         NA
2   NA 44.55    NA            NA   NA 28.76      64.32        130      38.86

Optional: Standardize Region Names Even More

If you want to map abbreviations (like Sw to South_West, S_S to South_South) to consistent, human-readable names, add a lookup step in the parse_string function:

# Define a lookup for consistent region names
region_lookup <- c(
  "Sw" = "South_West",
  "S_S" = "South_South",
  "North_Central" = "North_Central",
  "North_East" = "North_East",
  "SouthEast" = "South_East",
  "SouthWest" = "South_West"
)

# Update the region cleaning step in parse_string
regions <- matches[, 2] %>%
  str_trim() %>%
  str_to_title() %>%
  str_replace_all(c("\\s+" = "_", "\\-" = "_", "\\." = "_")) %>%
  recode(!!!region_lookup)

This ensures your column names are perfectly consistent, even if the input uses different abbreviations or spellings.

内容的提问来源于stack exchange,提问作者Tunde Awosanya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:46:01