You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中将特定格式的文本字符串转换为数据框?

Parsing Legislative String into Structured Data Frame

The core issue with using read.table() here is that it splits content purely on whitespace, which doesn't account for the meaningful grouping of data in parentheses. Instead, we can use regular expressions to target the specific patterns in your string, ensuring reliable extraction of each column even if names/districts have multiple components.

Step-by-Step Solution (R)

We'll use the stringr package (part of the tidyverse) for intuitive regex operations:

  • First, load the package and define your input string:
library(stringr)
headings_text <- "ADNOT (Philippe) sénateur (Aube) NI"
  • Extract each component using regex patterns tailored to your string's structure:

    • LastName: Capture all characters before the first opening parenthesis, then trim any extra whitespace.
    • FirstName: Extract the text inside the first pair of parentheses.
    • District: Extract the text inside the second pair of parentheses (after the title "sénateur").
    • Party: Capture the final token at the end of the string.
  • Combine these into a data frame with your desired column names:

parsed_data <- data.frame(
  LastName = str_trim(str_extract(headings_text, "^[^(]+")),
  FirstName = str_match(headings_text, "\\(([^)]+)\\)")[, 2],
  District = str_match_all(headings_text, "\\(([^)]+)\\)")[[1]][, 2][2],
  Party = str_extract(headings_text, "\\w+$"),
  stringsAsFactors = FALSE
)

Output

Running this code will give you the exact data frame you need:

LastName FirstName District Party
1    ADNOT  Philippe     Aube    NI

Robustness Notes

  • This approach works even if LastName has spaces (e.g., "DE LA FONTAINE (Jean) sénateur (Paris) LR" would correctly extract "DE LA FONTAINE" as LastName).
  • If you have a vector of such strings, the code will automatically vectorize—no need for loops.
  • If the input structure varies slightly (e.g., different titles), you can adjust the regex to ignore any text between the first and second parentheses (e.g., using str_replace to remove the title before extracting District).

内容的提问来源于stack exchange,提问作者dkro23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:16:06