You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中(优先用dplyr)基于指定关键词生成通顺随机句子

在R中用dplyr生成符合语法的随机中文句子

先把给定关键词按语法属性分组,再预设几种合规的中文句子模板,用dplyr配合随机抽样来生成指定数量的句子。

关键词分组(中文)

  • 名词:赫尔辛基、城镇、污染
  • 形容词:大、美好、很多
  • 限定词:一个
  • 副词:不
  • 动词:是

代码实现

library(dplyr)
library(glue)

# 构造关键词分组表
keywords <- tibble(
  category = c(rep("noun", 3), rep("adj", 3), "determiner", "adverb", "verb"),
  word = c("赫尔辛基", "城镇", "污染", "大", "美好", "很多", "一个", "不", "是")
)

# 单句生成函数
gen_sentence <- function() {
  # 预设语法通顺的模板
  templates <- list(
    list(tpl = "{adj} {noun}", groups = c("adj", "noun")),
    list(tpl = "{noun} {verb} {determiner} {adj} {noun}", groups = c("noun", "verb", "determiner", "adj", "noun")),
    list(tpl = "{adverb},{noun} {verb} {adj}", groups = c("adverb", "noun", "verb", "adj")),
    list(tpl = "{noun} {verb} {adj} {noun}", groups = c("noun", "verb", "adj", "noun"))
  )
  
  # 随机选模板
  chosen_tpl <- sample(templates, 1)[[1]]
  
  # 抽取对应关键词
  words <- purrr::map_chr(chosen_tpl$groups, ~{
    keywords %>% filter(category == .x) %>% pull(word) %>% sample(1)
  })
  
  # 替换模板生成句子
  glue(chosen_tpl$tpl, .envir = as.list(setNames(words, gsub("\\d", "", chosen_tpl$groups))))
}

# 生成n条句子(示例n=5)
n <- 5
result <- tibble(sentence = replicate(n, gen_sentence()))

print(result)

说明

  • 需要提前安装dplyr和glue包
  • 模板可根据需求自行调整,所有组合严格使用给定关键词,保证语法逻辑通顺
  • 通过分组抽样避免无意义的词组合,确保生成句子符合中文表达习惯

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 01:00:48