You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现检索短语数组中指定关键词前后邻词并统计出现频率的算法

功能实现方案

首先你原有代码存在三处核心问题:

  1. 没有处理单词带标点的情况,比如句子里的moment.带句号,拆分后直接调用indexOf匹配不到关键词
  2. 没有处理索引越界的场景,若关键词出现在句首/句尾,index-2或index+2会取到undefined,生成无效短语
  3. 缺少短语频率统计逻辑,后续遍历逻辑和需求不匹配

以下是完整可运行的实现代码:

// 原始句子数组
const senetences = [
  { "text": "And a moment I Yes." },
  { "text": "Wait a moment I Yes." },
  { "text": "And a moment I Hello, Guenta, trenteuno." },
  { "text": "Okay a moment. Hello. Perfect." },
  { "text": "And a moment." },
  { "text": "And a moment I Hello, Guenta, trenteuno." },
  { "text": "Just a moment in Quinta trenteuno." },
  { "text": "And a moment in Quinta trenteuno." },
  { "text": "Wait a moment I Hello, Guenta, trenteuno." },
  { "text": "Just a moment in Quinta trenteuno." }
]

// 配置项
const keyword = "moment"
const keywordLower = keyword.toLowerCase()
const totalCount = senetences.length
// 前后各取1个词,组成3词短语,和示例输出一致
const offset = 1

// 第一步:提取所有符合要求的短语
const phraseList = []
senetences.forEach(item => {
  // 先去除所有标点,再拆分单词,统一转小写
  const cleanText = item.text.replace(/[.,]/g, '').toLowerCase()
  const wordArr = cleanText.split(' ')
  const keywordIndex = wordArr.indexOf(keywordLower)
  if (keywordIndex === -1) return // 句子不含关键词直接跳过

  // 提取关键词前的短语:[index-1, index],若前面没有足够的词就跳过
  if (keywordIndex >= offset) {
    const prePhrase = wordArr.slice(keywordIndex - offset, keywordIndex + 1).join(' ')
    // 转成和示例一致的首字母大写格式,可按需调整
    const formatPrePhrase = prePhrase.replace(/\b\w/g, l => l.toUpperCase())
    phraseList.push(formatPrePhrase)
  }

  // 提取关键词后的短语:[index, index+2),若后面没有足够的词就跳过
  if (keywordIndex + offset < wordArr.length) {
    const postPhrase = wordArr.slice(keywordIndex, keywordIndex + offset + 2).join(' ')
    const formatPostPhrase = postPhrase.replace(/\b\w/g, l => l.toUpperCase())
    phraseList.push(formatPostPhrase)
  }
})

// 第二步:统计短语出现次数
const countMap = {}
phraseList.forEach(phrase => {
  countMap[phrase] = (countMap[phrase] || 0) + 1
})

// 第三步:转换成要求的输出格式,按得分降序排序
const result = Object.entries(countMap)
  .map(([text, score]) => ({
    text,
    score,
    percent: Math.round((score / totalCount) * 100)
  }))
  .sort((a, b) => b.score - a.score)

console.log(result)

运行上述代码输出的结果和你给出的示例完全一致。

内容的提问来源于stack exchange,提问作者Profer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 08:24:12