You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中使用正则表达式从文本数据框提取指定信息

Solution for Extracting Specific Details with Regex in R

Hi there! Let's work through how to pull those four key pieces of information from your text dataframe using regular expressions in R. We'll use the stringr package (part of the tidyverse) because it makes regex operations clean and intuitive.

Step 1: Setup and Your Original Data

First, let's confirm your input data and load the required package:

library(stringr)

# Your input dataframe
text <- data.frame(
  page = c(1,1,2,3),
  sen = c(1,2,1,1),
  text = c(
    "Dear Mr case 1",
    "the value of my property is £500,000.00 and it was built in 1980",
    "The protected percentage is 0% for 2 years",
    "The interest rate is fixed for 2 years at 4.8%"
  )
)

Step 2: Extract Each Target Field

We'll first combine all text into a single string to simplify extraction (since your info is spread across different rows). Then use regex to pull each specific detail:

# Combine all text entries into one string for easier cross-row extraction
full_text <- paste(text$text, collapse = " ")

# 1. Extract Gender (map "Mr" to Male)
gender <- ifelse(str_detect(full_text, "\\bMr\\b"), "Male", "Unknown")

# 2. Extract Property Value (matches standard pound currency format)
property_value <- str_extract(full_text, "£\\d{1,3}(,\\d{3})*(\\.\\d{2})?")

# 3. Extract Protected Percentage (targets the % value after the label phrase)
protected_percentage <- str_extract(full_text, "(?<=protected percentage is )\\d+%")

# 4. Extract Interest Rate (grabs the decimal rate after the context phrase)
interest_rate <- str_extract(full_text, "(?<=interest rate is fixed for \\d+ years at )\\d+\\.\\d+%")

Step 3: Build the Result Dataframe

Now assemble these extracted values into the structured dataframe you want:

result_df <- data.frame(
  Gender = gender,
  Property_Value = property_value,
  Protected_Percentage = protected_percentage,
  Interest_Rate = interest_rate,
  stringsAsFactors = FALSE
)

# View the final result
print(result_df)

What Each Regex Does

Let’s break down the patterns so you understand how they target the right data:

  • Gender: \\bMr\\b matches the whole word "Mr" (the \\b are word boundaries to avoid partial matches like "MrSmith"). We use str_detect to check for its presence and map to "Male".
  • Property Value: £\\d{1,3}(,\\d{3})*(\\.\\d{2})? matches pound symbols followed by standard currency formatting (supports thousands separators and decimal cents).
  • Protected Percentage: (?<=protected percentage is )\\d+% uses a positive lookbehind ((?<=...)) to only capture the number+% that comes right after the label phrase, ignoring the rest of the text.
  • Interest Rate: The lookbehind here targets the decimal rate specifically after the full context about fixed years, ensuring we don't accidentally pick up other percentage values.

Quick Alternative: Extract Directly from Rows

If you prefer to pull each field directly from its specific row (without combining text), you can target individual entries in the text column:

# Extract property value from the second row directly
property_value <- str_extract(text$text[2], "£\\d{1,3}(,\\d{3})*(\\.\\d{2})?")

内容的提问来源于stack exchange,提问作者SAJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:57:16