正则截取最多20个单词时末尾重复词问题原因咨询
问题背景
要通过正则截取字符串,最多保留20个单词,待处理文本如下:
const text = `Space travel is the ultimate adventure! Imagine soaring past the stars and exploring new worlds. It's the stuff of dreams and science fiction, but believe it or not, space travel is a real thing. Humans and robots are constantly venturing out into the cosmos to uncover its secrets and push the boundaries of what's possible`;
使用的正则模式:
const regex = new RegExp(`^(\\S+\\s+){1,20}`);
执行text.match(regex)时,预期得到:
Space travel is the ultimate adventure! Imagine soaring past the stars
and exploring new worlds. It's the stuff of dreams
但实际结果末尾多了重复的“dreams”:
Space travel is the ultimate adventure! Imagine soaring past the stars
and exploring new worlds. It's the stuff of dreams dreams
用regex.exec(text)返回的数组包含2个元素,第二个元素就是这个重复词。添加g标志或把(\S+\s+)改为非捕获组^(?:\S+\s+){1,20}即可解决问题,下面解释原因。
核心原因
1. 捕获组的重复匹配特性
你用的正则里,(\S+\s+)是一个捕获组,并且被设定为重复1到20次。正则引擎处理重复的捕获组时,每次匹配都会覆盖捕获组的存储内容,最终捕获组里只会保留最后一次匹配到的内容。在你的场景里,最后一次匹配到的\S+\s+就是dreams (“dreams”加上后面的空格)。
2. match()方法的返回规则
当正则没有g修饰符时,String.prototype.match()返回的数组结构是:[整个正则匹配的完整结果, 捕获组1的内容, 捕获组2的内容, ...]。所以你的场景里,match()返回的数组是:
[ "Space travel is the ultimate adventure! Imagine soaring past the stars\nand exploring new worlds. It's the stuff of dreams ", "dreams " ]
如果直接输出这个数组(比如在控制台打印,或者不小心拼接了数组元素),就会看起来像是末尾多了一个“dreams”(空格在显示中通常被忽略)。
3. 两种解决方法的原理
- 非捕获组
(?:\S+\s+):非捕获组的作用是分组匹配,但不会保存匹配的内容。使用它之后,match()返回的数组就只有整个匹配结果这一个元素,不会出现额外的捕获内容。 - 添加
g修饰符:当正则带有g修饰符时,match()会只返回所有完整匹配的结果数组,不会包含捕获组的内容,因此也只会得到你预期的前20个单词组合。
内容的提问来源于stack exchange,提问作者Vsevolod IV

