You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从文件指定位置提取文本并构建含编号、称谓、姓名、内容的DataFrame?

Extract Structured "Who Said What" Data with R

Hey there! Let's turn that unstructured text into the clean DataFrame you need—capturing the number label, title (Mr./Mrs.), name, and content for each speech. I'll use R with stringr for regex magic and dplyr for data wrangling, which fits your requirements perfectly.

Step 1: Prep Your Raw Data

First, let's clean up the original text (remove those extra ** markers) and load the necessary packages:

library(stringr)
library(dplyr)

# Your raw text (cleaned of bold markers)
text <- "9 Mr.ABCD. Content1. Mrs. DEFG.Content2. 8 Mr.DBC something else. Content3."

Step 2: Split Text into Numbered Blocks

We'll split the text into chunks where each chunk starts with a number (like 9 or 8)—each chunk holds all speeches under that number:

# Extract all blocks starting with a number
content_blocks <- str_extract_all(text, "\\d+.*?(?=\\d|$)")[[1]]

This regex grabs everything from a number up until the next number or the end of the text, giving us two blocks: one for 9, one for 8.

Step 3: Break Down Blocks into Individual Speeches & Extract Fields

Now we'll loop through each block, pull out the number, split the block into individual speeches, and use regex to pull out the title, name, and content:

# Empty list to store our results
result_list <- list()

for (block in content_blocks) {
  # Grab the number (convert to integer for clean typing)
  number <- str_extract(block, "\\d+") %>% as.integer()
  
  # Split the block into separate speeches (split when we see ". " followed by Mr./Mrs.)
  speeches <- str_split(block, "(?<=\\.) (?=(Mr\\.|Mrs\\.))")[[1]]
  # Remove the number prefix from the first speech in the block
  speeches[1] <- str_remove(speeches[1], "^\\d+ ")
  
  # Process each speech
  for (speech in speeches) {
    # Regex to match: Title (Mr./Mrs.) → All-uppercase Name → Content after the period
    match <- str_match(speech, "(Mr\\.|Mrs\\.) ([A-Z]+)\\.?\\s*(.*)")
    
    # Turn the match into a tiny data frame row
    row_df <- tibble(
      number = number,
      title = match[2],
      name = match[3],
      content = str_trim(str_remove(match[4], "\\.$"))  # Trim spaces and remove trailing period
    )
    
    # Add the row to our result list
    result_list <- append(result_list, list(row_df))
  }
}

# Combine all rows into the final DataFrame
final_df <- bind_rows(result_list)

Step 4: Check Your Final DataFrame

When you run print(final_df), you'll get exactly the structure you wanted:

# # A tibble: 3 × 4
#   number title name  content          
#    <int> <chr> <chr> <chr>            
# 1      9 Mr.   ABCD  Content1        
# 2      9 Mrs.  DEFG  Content2        
# 3      8 Mr.   DBC   something else. Content3

Quick Regex Breakdown

  • \\d+.*?(?=\\d|$): Grabs numbered blocks by matching digits, then everything until the next digit or end of text.
  • (?<=\\.) (?=(Mr\\.|Mrs\\.)): Splits speeches by looking for a period followed by a space and then Mr./Mrs. (no messy splitting inside content!).
  • (Mr\\.|Mrs\\.) ([A-Z]+)\\.?\\s*(.*): Targets your exact rules—grabs the title, then the all-uppercase name, then everything after as content.

This solution fits your core need perfectly: turning unlabeled text into structured data showing who said what, with their number tag attached.

内容的提问来源于stack exchange,提问作者Foulball

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 14:02:35