如何使用可全匹配子串的正则拆分字符串并校验无无效字符
实现方案
核心思路是跟踪每次匹配的位置,校验匹配结果之间是否存在未被正则覆盖的无效片段,具体实现代码如下:
var re = /\s*(\w+|"[^"]*")\s*/g // 匹配单词或带双引号的字符串 var str = 'hey ??? "a"b' // '???' 片段不属于有效token let lastEndPos = 0 const validTokens = [] for (const match of str.matchAll(re)) { // 校验当前匹配起始位置和上一次匹配结束位置之间是否存在无效内容 if (match.index > lastEndPos) { const invalidSlice = str.slice(lastEndPos, match.index) throw new Error(`检测到无效片段:${invalidSlice},起始位置:${lastEndPos}`) } validTokens.push(match[1]) // 更新上一次匹配的结束位置 lastEndPos = match.index + match[0].length } // 校验字符串末尾是否残留未匹配的无效内容 if (lastEndPos < str.length) { const invalidSlice = str.slice(lastEndPos) throw new Error(`检测到末尾无效片段:${invalidSlice},起始位置:${lastEndPos}`) } // 无异常时输出结果 validTokens.forEach(token => console.log("Matched:", token))
方案说明
- 每次匹配后记录当前匹配的结束位置,下一次匹配开始时先校验起始位置是否和上一次结束位置对齐,存在差值就说明中间有未匹配的无效片段
- 所有匹配完成后还要校验末尾是否有未覆盖的内容,避免末尾残留无效字符的场景漏检
- 你给出的测试用例运行该代码会直接抛出错误,提示
???为无效片段,符合预期要求
内容的提问来源于stack exchange,提问作者Blue Nebula
相关产品推荐
相关产品推荐

