在R语言中使用正则表达式从文本数据框提取指定信息
Hi there! Let's work through how to pull those four key pieces of information from your text dataframe using regular expressions in R. We'll use the stringr package (part of the tidyverse) because it makes regex operations clean and intuitive.
Step 1: Setup and Your Original Data
First, let's confirm your input data and load the required package:
library(stringr) # Your input dataframe text <- data.frame( page = c(1,1,2,3), sen = c(1,2,1,1), text = c( "Dear Mr case 1", "the value of my property is £500,000.00 and it was built in 1980", "The protected percentage is 0% for 2 years", "The interest rate is fixed for 2 years at 4.8%" ) )
Step 2: Extract Each Target Field
We'll first combine all text into a single string to simplify extraction (since your info is spread across different rows). Then use regex to pull each specific detail:
# Combine all text entries into one string for easier cross-row extraction full_text <- paste(text$text, collapse = " ") # 1. Extract Gender (map "Mr" to Male) gender <- ifelse(str_detect(full_text, "\\bMr\\b"), "Male", "Unknown") # 2. Extract Property Value (matches standard pound currency format) property_value <- str_extract(full_text, "£\\d{1,3}(,\\d{3})*(\\.\\d{2})?") # 3. Extract Protected Percentage (targets the % value after the label phrase) protected_percentage <- str_extract(full_text, "(?<=protected percentage is )\\d+%") # 4. Extract Interest Rate (grabs the decimal rate after the context phrase) interest_rate <- str_extract(full_text, "(?<=interest rate is fixed for \\d+ years at )\\d+\\.\\d+%")
Step 3: Build the Result Dataframe
Now assemble these extracted values into the structured dataframe you want:
result_df <- data.frame( Gender = gender, Property_Value = property_value, Protected_Percentage = protected_percentage, Interest_Rate = interest_rate, stringsAsFactors = FALSE ) # View the final result print(result_df)
What Each Regex Does
Let’s break down the patterns so you understand how they target the right data:
- Gender:
\\bMr\\bmatches the whole word "Mr" (the\\bare word boundaries to avoid partial matches like "MrSmith"). We usestr_detectto check for its presence and map to "Male". - Property Value:
£\\d{1,3}(,\\d{3})*(\\.\\d{2})?matches pound symbols followed by standard currency formatting (supports thousands separators and decimal cents). - Protected Percentage:
(?<=protected percentage is )\\d+%uses a positive lookbehind ((?<=...)) to only capture the number+% that comes right after the label phrase, ignoring the rest of the text. - Interest Rate: The lookbehind here targets the decimal rate specifically after the full context about fixed years, ensuring we don't accidentally pick up other percentage values.
Quick Alternative: Extract Directly from Rows
If you prefer to pull each field directly from its specific row (without combining text), you can target individual entries in the text column:
# Extract property value from the second row directly property_value <- str_extract(text$text[2], "£\\d{1,3}(,\\d{3})*(\\.\\d{2})?")
内容的提问来源于stack exchange,提问作者SAJ

