如何使用正则实现含匹配项的指定字数文本片段截取?
实现文本片段截取函数
getSnippet 核心思路
单纯用正则很难灵活处理「一侧单词不足时从另一侧补足」的逻辑,更可靠的方案是先将文本拆分为单词数组,再针对每个匹配项计算片段范围:
- 把文本按空白字符拆分为单词数组(单词包含附带的标点)
- 找出所有不区分大小写匹配目标字符串的单词索引
- 对每个匹配索引,计算片段的起始/结束位置:先按总数平分前后单词数,若某一侧数量不足,剩余数量从另一侧补充
- 拼接对应范围的单词为字符串,返回结果集合
代码实现(JavaScript)
function getSnippet(text, number, match) { // 拆分文本为单词数组(按任意空白字符分割) const words = text.split(/\s+/); const target = match.toLowerCase(); // 收集所有匹配目标的单词索引(不区分大小写) const matchPositions = words.reduce((acc, word, index) => { if (word.toLowerCase().includes(target)) { acc.push(index); } return acc; }, []); return matchPositions.map(pos => { const totalWords = words.length; const extraCount = number - 1; // 匹配项之外需要的单词总数 let leftWanted = Math.floor(extraCount / 2); let rightWanted = extraCount - leftWanted; // 计算实际能取到的左侧单词数 let leftActual = Math.min(leftWanted, pos); // 剩余需要从右侧取的数量 let rightNeeded = extraCount - leftActual; // 实际能取到的右侧单词数 let rightActual = Math.min(rightNeeded, totalWords - pos - 1); // 若右侧不足,从左侧补足 if (rightActual < rightNeeded) { const leftNeeded = extraCount - rightActual; leftActual = Math.min(leftNeeded, pos); } const startIdx = pos - leftActual; const endIdx = pos + rightActual; // 拼接单词为片段字符串 return words.slice(startIdx, endIdx + 1).join(' '); }); } // 测试示例 const sampleText = "Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s"; console.log(getSnippet(sampleText, 10, 'ipsum')); // 输出:["Lorem Ipsum is simply dummy text of the printing and", "and typesetting industry. Lorem Ipsum has been the industry's standard"]
代码说明
- 单词拆分:用
/\s+/正则分割,兼容多个连续空白字符的情况 - 匹配逻辑:统一转为小写判断,实现不区分大小写的模糊匹配
- 范围计算:优先按总数平分前后单词,若某侧单词不足,自动从另一侧补充,确保每个片段的单词总数严格等于
number
内容的提问来源于stack exchange,提问作者Ncifra
相关产品推荐
相关产品推荐

