如何按给定规则切割指定格式字符串并转换为目标结构化数组
实现方案
核心逻辑是基于关键词匹配位置做文本切片,交替插入普通文本段和替换后的关键词配置段,具体执行流程:
- 先对所有关键词做长度降序排序,防止短关键词提前匹配,截断更长的包含型关键词
- 从文本起始位置开始遍历,每次查找当前游标之后最先出现的匹配关键词
- 若游标位置到关键词起始位置之间存在有效文本,将这段文本作为纯文本项(仅含
text字段)存入结果 - 匹配到的关键词不保留原文本,替换为配置中指定的
replaceKeyword,同时携带配置里的link、color属性作为特殊项存入结果 - 将游标移动到当前匹配关键词的结束位置,重复上述查找流程,直到游标走到文本末尾
- 遍历结束后如果游标后还有剩余文本,补为纯文本项,最后过滤掉
text为空的无效项即可
代码实现(JavaScript)
function splitContentByKeywords(input) { const { content, keywords } = input; // 关键词按长度降序排列,避免短词优先匹配截断长词 const sortedKeywords = [...keywords].sort((a, b) => b.keyword.length - a.keyword.length); const result = []; let currentIndex = 0; const contentLength = content.length; while (currentIndex < contentLength) { let earliestMatch = null; // 查找当前游标之后最早出现的关键词 for (const kwConfig of sortedKeywords) { const matchIndex = content.indexOf(kwConfig.keyword, currentIndex); if (matchIndex === -1) continue; if (!earliestMatch || matchIndex < earliestMatch.index) { earliestMatch = { index: matchIndex, config: kwConfig, endIndex: matchIndex + kwConfig.keyword.length }; } } if (!earliestMatch) { // 无剩余匹配,将后续所有文本存为普通项 const restText = content.slice(currentIndex); if (restText) result.push({ text: restText }); break; } // 存入匹配段之前的普通文本 const normalText = content.slice(currentIndex, earliestMatch.index); if (normalText) result.push({ text: normalText }); // 存入替换后的关键词项 result.push({ text: earliestMatch.config.replaceKeyword, link: earliestMatch.config.link, color: earliestMatch.config.color }); // 移动游标到当前匹配段末尾 currentIndex = earliestMatch.endIndex; } return result; } // 调用示例 const inputData = { content: 'key1vip has been serving you for 2 days, and the customer service will provide you with professional answers and formulate solutions', keywords: [{ keyword: 'key1', replaceKeyword: 'On', link: '', color: '', }, { keyword: '2', // 原示例此处笔误写为key2,实际要匹配原文中的数字2才能替换为30 replaceKeyword: '30', link: '', color: '', }] }; const _array = splitContentByKeywords(inputData);
注意事项
- 配置关键词时需要和待匹配的原文内容完全一致,否则无法命中匹配规则
- 如果业务场景需要忽略大小写匹配,可以在匹配前统一把原文和关键词转为相同大小写再做位置查找
- 若关键词之间存在互相包含的关系,必须保留长度降序排序的逻辑,否则会出现匹配错误
- 不需要逐字符移动游标,每次匹配完成后直接跳到匹配段末尾即可,能提升长文本下的处理效率
内容的提问来源于stack exchange,提问作者Nemo
相关产品推荐
相关产品推荐

