You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R实现类似Excel快速填充的动态问卷字符串品牌名提取功能

R语言实现问卷品牌名批量提取(对标Excel快速填充功能)

需求背景

处理问卷数据集时,需要从句式不固定的问题文本中提取品牌名,原始数据示例如下:

Data #1
What do you know about AlphaToy?

Data #2
What comes to your mind when you heard AlphaCars?

Data #3
What do you think of FoodTruckers?

需要提取AlphaToy、AlphaCars、FoodTruckers这类品牌名,该需求在Excel中可通过快速填充(flash fill)实现,演示动图如下:
Excel快速填充提取品牌名演示

期望实现效果

要求用R语言封装类似快速填充的功能函数,适配动态字符串,输入输出要求如下:

测试样例数据

brandName <- list(
  Toy = c(
    "1. What do you know about AlphaToy?",
    "2. What do you know about BetaToyz?",
    "3. What do you know about CharlieDoll?",
    "4. What do you know about DeltaToys?",
    "5. What do you know about Echoty?"
  ),
  Car = c(
    "18. What comes to your mind when you heard AlphaCars?",
    "19. What comes to your mind when you heard BestCar?",
    "20. What comes to your mind when you heard CoolCarz?"
  ),
  Trucker = c(
    "5. What do you think of FoodTruckers?",
    "6. What do you think of IceCreamTruckers?",
    "7. What do you think of JellyTruckers?",
    "8. What do you think of SodaTruckers?"
  )
)

extractBrandName <- function(...) {
  # 待实现代码
}

单组输入期望输出

> extractBrandName(brandName$Toy)
[1] "AlphaToy"    "BetaToyz"    "CharlieDoll" "DeltaToys"   "Echoty"

批量处理期望输出

> lapply(brandName, extractBrandName)
$Toy
[1] "AlphaToy"    "BetaToyz"    "CharlieDoll" "DeltaToys"   "Echoty"     

$Car
[1] "AlphaCars" "BestCar"   "CoolCarz" 

$Trucker
[1] "FoodTruckers"     "IceCreamTruckers" "JellyTruckers"    "SodaTruckers"

适配规则说明

  • 品牌名可能为小写、大写格式,也可能由2个及以上单词组成,例如IBM、Louis Vuitton
  • 品牌名可能出现在句子任意位置,不固定在句尾,不同客户提供的问卷句式存在差异,无法提前预判

已实现方案

实现思路

核心逻辑为提取输入字符串的公共词,过滤公共词后剩余内容即为品牌名:通过Reduce()包裹intersect()获取所有字符串的公共词,再通过lapply()过滤公共词,使用str_c(collapse = " ")拼接多词组成的品牌名。

代码实现

library(stringr)

extractBrandName <- function(x) {
  cleanWords <- x %>%
    str_remove_all("^\\d+|\\.|,|\\?") %>% 
    str_squish() %>% 
    str_split(" ")
  commonWords <- cleanWords %>% 
    Reduce(intersect, .)
  extractedWords <- cleanWords %>% 
    lapply(., function(y) {
      y[!y %in% commonWords] %>% 
        str_c(collapse = " ")
    }) %>% unlist()
  return(extractedWords)
}

测试效果

测试用例1(基础场景)

> extractBrandName(brandName$Toy)
[1] "AlphaToy"    "BetaToyz"    "CharlieDoll" "DeltaToys"   "Echoty"     
> lapply(brandName, extractBrandName)
$Toy
[1] "AlphaToy"    "BetaToyz"    "CharlieDoll" "DeltaToys"   "Echoty"     

$Car
[1] "AlphaCars" "BestCar"   "CoolCarz" 

$Trucker
[1] "FoodTruckers"     "IceCreamTruckers" "JellyTruckers"    "SodaTruckers"    

测试用例2(复杂场景:多词品牌、品牌位置不固定)

测试数据:

brandName2 <- list(
  Middle = c("Have you used any products from AlphaToy this past 6 months?",
             "Have you used any products from BetaToys Collection this past 6 months?",
             "Have you used any products from Charl TOYZ this past 6 months?"),
  First = c("AlphaCars is the best automobile dealer, yes/no?",
            "Best Vehc is the best automobile dealer, yes/no?",
            "CoolCarz & Bike is the best automobile dealer, yes/no?")
)

输出结果:

> lapply(brandName2, extractBrandName)
$Middle
[1] "AlphaToy"            "BetaToys Collection" "Charl TOYZ"         

$First
[1] "AlphaCars"       "Best Vehc"       "CoolCarz & Bike"

内容的提问来源于stack exchange,提问作者rifset

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 08:36:04