You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中提取扫描文档文本中的全部姓名(含亲属姓名)

Extracting Names from Scanned Text in R

Hey there! Let's figure out how to pull all the names (personal and parent/husband/mother) from your scanned document text in R. The text has some messy formatting, but regular expressions will help us target exactly what we need.

Step 1: Load the Sample Text

First, let's store your scanned text in an R variable so we can work with it:

# Paste your scanned text here
text_data <- "Name ; Nagesh V Name ; Savitha Name ; A T Vitalrao
Father's Le Venkatappa Father's Srinivas Father's Thirumagandam
Name: Name: Name:

House No.:9/1 House No.:9/C House No.:9/C

Age: 60 Sex: Male Age: 28 Sex: Female Age: 85 Sex: Male
BCW1799964 BCW1797224 SOH0004515
Name : V Kedarnath Name : K Nalini Name : Sayiraj
Fa..."

Step 2: Extract Personal Names

Your text uses a few variations for personal names: Name ; , Name : , and even Name: (with no space). We'll use the stringr package (part of the tidyverse) to match all these patterns with a regular expression:

# Install the package first if you haven't: install.packages("stringr")
library(stringr)

# Extract all personal names using a regex that handles different separators
personal_names <- str_extract_all(text_data, "(?<=Name\\s*(;|:)\\s)[A-Za-z\\s]+")[[1]]

# Clean up extra whitespace and remove empty entries (from the "Name: Name: Name:" line)
personal_names <- str_trim(personal_names)
personal_names <- personal_names[personal_names != ""]

Running this will give you a vector of personal names like:
"Nagesh V" "Savitha" "A T Vitalrao" "V Kedarnath" "K Nalini" "Sayiraj"

Step 3: Extract Parent/Husband/Mother Names

For names prefixed with Father's, we can use another regex to target those specifically. If your text also includes Mother's or Husband's, we can easily adjust the pattern to include those too:

# Extract names from Father's entries (adjust the regex to add Mother's/Husband's if needed)
parent_names <- str_extract_all(text_data, "(?<=Father's\\s)[A-Za-z\\s]+")[[1]]

# Clean up whitespace
parent_names <- str_trim(parent_names)

This will return:
"Le Venkatappa" "Srinivas" "Thirumagandam"

Step 4: Combine All Names (Optional)

If you want a single list of all unique names, combine the two vectors and remove duplicates:

all_names <- unique(c(personal_names, parent_names))

Notes for Edge Cases

  • If your full text has other prefixes like Mother's or Husband's, update the parent name regex to:
    parent_names <- str_extract_all(text_data, "(?<=(Father's|Mother's|Husband's)\\s)[A-Za-z\\s]+")[[1]]
    
  • If some names include initials or hyphens, tweak the regex character set to [A-Za-z\\s\\.-] to capture those.

内容的提问来源于stack exchange,提问作者Sharaf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:38:05