You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中提取Date:、City:、Note:后的对应内容?

问题描述

现有如下CSV格式的表格数据:

No, Memo
  1, Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.
  2, Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.
  3, Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.
  4, Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely.

需要从Memo列中提取三个字段:

  • Date:与City:之间的日期内容
  • City:与Note:之间的城市代码
  • Note:之后的完整语句

最终期望输出格式如下:

No Date       City Note
  1 2020/10/22 UA   True mastery of any skill takes a lifetime.
  2 2022/11/01 CH   Sweat is the lubricant of success.
  3 2022y11m1d UA   Every noble work is at first impossible.
  4 2022y2m15d AA   Live beautifully, dream passionately, love completely.
可行实现方法

方法1:使用tidyr::extract(简洁高效)

借助tidyr包的extract函数,通过正则表达式直接拆分Memo列,一步完成字段提取:

library(tidyr)
library(readr)

# 读取原始CSV数据(trim_ws = TRUE 自动去除单元格前后空格)
df <- read_csv("""
No, Memo
  1, Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.
  2, Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.
  3, Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.
  4, Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely.
""", trim_ws = TRUE)

# 拆分Memo列,提取目标字段
df_clean <- df %>%
  extract(
    Memo, 
    into = c("Date", "City", "Note"),
    regex = "Date: (.*?) City: (.*?) Note: (.*)",
    remove = FALSE # 若无需保留原Memo列,设为TRUE即可
  )

# 输出整理后的结果
print(df_clean[, c("No", "Date", "City", "Note")])

正则说明

  • Date: (.*?):非贪婪匹配Date:后到 City:前的所有内容
  • City: (.*?):非贪婪匹配City:后到 Note:前的所有内容
  • Note: (.*):贪婪匹配Note:后的所有剩余内容(因为是最后一段,无需担心过度匹配)

方法2:使用stringr包的正则提取

通过stringr提供的str_extract函数,结合正则断言精准提取每个字段:

library(stringr)

# 构造数据框
df <- data.frame(
  No = 1:4,
  Memo = c(
    "Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.",
    "Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.",
    "Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.",
    "Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely."
  ),
  stringsAsFactors = FALSE
)

# 逐个提取字段
df$Date <- str_extract(df$Memo, "(?<=Date: ).*?(?= City:)")
df$City <- str_extract(df$Memo, "(?<=City: ).*?(?= Note:)")
df$Note <- str_extract(df$Memo, "(?<=Note: ).*")

# 整理并输出结果
result <- df[, c("No", "Date", "City", "Note")]
print(result)

正则断言说明

  • (?<=Date: ):正向后行断言,匹配前面是Date: 的位置
  • .*?(?= City:):正向先行断言,匹配到 City:前的内容(非贪婪模式避免过度匹配)
  • (?<=Note: ).*:匹配Note: 后的所有剩余内容

方法3:Base R原生实现(无需额外包)

使用Base R的regexec和regmatches函数,不依赖第三方包完成提取:

# 构造数据框
df <- data.frame(
  No = 1:4,
  Memo = c(
    "Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.",
    "Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.",
    "Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.",
    "Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely."
  ),
  stringsAsFactors = FALSE
)

# 定义匹配正则
pattern <- "Date: (.*?) City: (.*?) Note: (.*)"

# 提取所有匹配结果
matches <- regmatches(df$Memo, regexec(pattern, df$Memo))

# 将匹配结果转换为数据框并合并
extracted <- do.call(rbind, lapply(matches, function(x) x[-1]))
colnames(extracted) <- c("Date", "City", "Note")
result <- cbind(df["No"], extracted)

# 输出结果
print(result)

逻辑说明

  • regexec:返回每个字符串中匹配正则的位置信息
  • regmatches:根据位置信息提取对应内容
  • 通过lapply和do.call(rbind, ...)将列表格式的匹配结果转换为数据框,最后与原No列合并

内容的提问来源于stack exchange,提问作者mashimena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 14:54:53