正则捕获尾部及行首星号但排除包裹强调文本星号的问题排查
问题原因
你编写的负向后行断言逻辑存在局限性:原断言(?<!\W\*\w([^*]|\w\*)*?)限定了星号后必须紧跟单字符\w,仅能匹配单字符斜体*a*的场景,当斜体内容为多字符、包含空格时,该断言无法命中,导致对应闭合星号没有被排除,会被误识别为脚注星号替换。
解决方案
将正则中的负向后行断言替换为更通用的逻辑,判断当前星号不是斜体的闭合标记即可:
const regex = /(?<=\S)(?<!\*[^*]*)\*+(?=[^\w*]|$)|^\*+(?=\s\S)/gm;
核心修改点:用(?<!\*[^*]*)替代原有的复杂断言,逻辑为「当前星号前面不存在单个*加非星号内容的组合」,直接排除所有斜体的闭合右星号,适配任意长度、包含空格的斜体内容。
完整修正后代码
/** * (?<=\S) -- 匹配前缀为非空白字符 * (?<!\*[^*]*) -- 排除斜体闭合星号:前面不存在未配对的*加非星号内容 * \*+ -- 匹配1个及以上星号(脚注标记主体) * (?=[^\w*]|$) -- 匹配后缀为非单词非星号字符,或行尾 * | -- 或 * ^\*+(?=\s\S) -- 匹配行首的星号,后接空格加非空白字符 */ const regex = /(?<=\S)(?<!\*[^*]*)\*+(?=[^\w*]|$)|^\*+(?=\s\S)/gm; const transform = m => { const superTable = [ '⁰', '¹', '²', '³', '⁴', '⁵', '⁶', '⁷', '⁸', '⁹' ]; let str = []; for (let len = m.length; len; len = (len - len % 10) / 10) { str.unshift(superTable[len % 10]); } return str.join(''); } /** [input, expectedOutput] */ const testCases = [ [`A b*** c`, `A b³ c`], [`A *b* c*`, `A *b* c¹`], [`A *b* *c* d*`, `A *b* *c* d¹`], [`A *b* c* d**`, `A *b* c¹ d²`], [`** a b c`, `² a b c`], [`** a b*** c`, `² a b³ c`], [`A *bc* d**`, `A *bc* d²`], [`A *b c* d**`, `A *b c* d²`], ]; const results = ['Input\t\t=>\tActual\t\t===\tExpected\t: Success']; results.push('='.repeat(73)); for (const [input, expected] of testCases) { const actual = input.replace(regex, transform); const extraSpacing = actual.length < 8 ? '\t' : ''; const success = actual === expected; results.push(`${input}\t=>\t${actual}${extraSpacing}\t===\t${expected}${extraSpacing}\t: ${success}`); } console.log(results.join('\n'));
运行后所有测试用例均返回success,符合预期。
内容的提问来源于stack exchange,提问作者dx_over_dt
相关产品推荐
相关产品推荐

