You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用gsub删除文本Customer行及多余空行的实现方法

解决方案

完全可以通过单个gsub函数配合Perl正则表达式同时实现删除Customer行和多余开头空行的需求。

核心实现代码

首先先模拟你的测试数据框:

# 构造示例数据
df <- data.frame(
  raw_text = c(
    "\"***ORDER LIST***\nCustomer: Lucille\nitem1: apples\nitem2: oranges\"",
    "\"***ORDER LIST***\nCustomer: Frank and Sally\nitem1: wine\nitem2: milk\"",
    "\"***ORDER LIST***\n\n\nitem1: wine\nitem2: milk\""
  ),
  stringsAsFactors = FALSE
)

# 执行文本清洗
df$cleaned_text <- gsub(
  pattern = "(\\*\\*\\*ORDER LIST\\*\\*\\*\\R)(?:Customer:[^\r\n]*\\R|\\R+)",
  replacement = "\\1",
  x = df$raw_text,
  perl = TRUE
)

正则规则说明

正则表达式分两部分逻辑:

  • 捕获分组(\\*\\*\\*ORDER LIST\\*\\*\\*\\R):匹配固定开头的订单头+紧跟的第一个换行符,替换时会完整保留这部分内容
  • 非捕获分组(?:Customer:[^\r\n]*\\R|\\R+):匹配需要删除的内容,两种匹配规则满足其一即可:
    • Customer:[^\r\n]*\\R:匹配整行以Customer:开头的内容,包含行尾的换行符
    • \\R+:匹配多个连续的换行符(也就是订单头后多余的空行)
  • 替换时仅保留捕获的订单头内容,即可同时完成两类冗余内容的删除

结果验证

清洗后的cleaned_text列内容和你给出的预期结果完全一致:

cat(df$cleaned_text[1])
# 输出:
# "***ORDER LIST***
# item1: apples
# item2: oranges"

内容的提问来源于stack exchange,提问作者Benjamin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 04:48:05