You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按给定规则切割指定格式字符串并转换为目标结构化数组

实现方案

核心逻辑是基于关键词匹配位置做文本切片,交替插入普通文本段和替换后的关键词配置段,具体执行流程:

  • 先对所有关键词做长度降序排序,防止短关键词提前匹配,截断更长的包含型关键词
  • 从文本起始位置开始遍历,每次查找当前游标之后最先出现的匹配关键词
  • 若游标位置到关键词起始位置之间存在有效文本,将这段文本作为纯文本项(仅含text字段)存入结果
  • 匹配到的关键词不保留原文本,替换为配置中指定的replaceKeyword,同时携带配置里的link、color属性作为特殊项存入结果
  • 将游标移动到当前匹配关键词的结束位置,重复上述查找流程,直到游标走到文本末尾
  • 遍历结束后如果游标后还有剩余文本,补为纯文本项,最后过滤掉text为空的无效项即可
代码实现(JavaScript)
function splitContentByKeywords(input) {
  const { content, keywords } = input;
  // 关键词按长度降序排列,避免短词优先匹配截断长词
  const sortedKeywords = [...keywords].sort((a, b) => b.keyword.length - a.keyword.length);
  const result = [];
  let currentIndex = 0;
  const contentLength = content.length;

  while (currentIndex < contentLength) {
    let earliestMatch = null;
    // 查找当前游标之后最早出现的关键词
    for (const kwConfig of sortedKeywords) {
      const matchIndex = content.indexOf(kwConfig.keyword, currentIndex);
      if (matchIndex === -1) continue;
      if (!earliestMatch || matchIndex < earliestMatch.index) {
        earliestMatch = {
          index: matchIndex,
          config: kwConfig,
          endIndex: matchIndex + kwConfig.keyword.length
        };
      }
    }

    if (!earliestMatch) {
      // 无剩余匹配,将后续所有文本存为普通项
      const restText = content.slice(currentIndex);
      if (restText) result.push({ text: restText });
      break;
    }

    // 存入匹配段之前的普通文本
    const normalText = content.slice(currentIndex, earliestMatch.index);
    if (normalText) result.push({ text: normalText });
    // 存入替换后的关键词项
    result.push({
      text: earliestMatch.config.replaceKeyword,
      link: earliestMatch.config.link,
      color: earliestMatch.config.color
    });
    // 移动游标到当前匹配段末尾
    currentIndex = earliestMatch.endIndex;
  }

  return result;
}

// 调用示例
const inputData = {
  content: 'key1vip has been serving you for 2 days, and the customer service will provide you with professional answers and formulate solutions',
  keywords: [{
    keyword: 'key1',
    replaceKeyword: 'On',
    link: '',
    color: '',
  }, {
    keyword: '2', // 原示例此处笔误写为key2,实际要匹配原文中的数字2才能替换为30
    replaceKeyword: '30',
    link: '',
    color: '',
  }]
};

const _array = splitContentByKeywords(inputData);
注意事项
  • 配置关键词时需要和待匹配的原文内容完全一致,否则无法命中匹配规则
  • 如果业务场景需要忽略大小写匹配,可以在匹配前统一把原文和关键词转为相同大小写再做位置查找
  • 若关键词之间存在互相包含的关系,必须保留长度降序排序的逻辑,否则会出现匹配错误
  • 不需要逐字符移动游标,每次匹配完成后直接跳到匹配段末尾即可,能提升长文本下的处理效率

内容的提问来源于stack exchange,提问作者Nemo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 21:39:04