正则表达式需求:提取4个连续大写字母后的全部文本
解决方案:匹配任意4个连续大写字母后的所有文本
嘿,刚好碰到过类似的需求,给你一个灵活度拉满的正则方案——不管是哪四个连续大写字母(不管是ABCD还是LOND这种),都能精准提取从包含它们的那一行开始的所有后续文本:
^.*?([A-Z]{4}).*$([\s\S]*)
正则工作原理拆解:
^.*?([A-Z]{4}).*$:用非贪婪的.*?定位第一个出现4个连续大写字母的整行,确保不会跳过更早的匹配项([\s\S]*):捕获这一行之后的所有内容,[\s\S]能匹配包括换行符在内的任意字符,*直接把剩下的所有文本都捞进来
实际示例演示
假设你的输入文本是:
Some warm-up text here.
No caps in this line either.
LONDON, UK. January 28, 2019-- More example of text, lots of text, Text text.
Imagine this is a long article... blah blah blah blah blah blah
Even more lines to test with.
用这个正则匹配后,你会得到想要的结果:
LONDON, UK. January 28, 2019-- More example of text, lots of text, Text text. Imagine this is a long article... blah blah blah blah blah blah Even more lines to test with.
如果你的正则引擎支持跨行正向预查,还可以用更简洁的版本:
(?<=^.*[A-Z]{4}.*$)[\s\S]*
不过要注意,部分老版本的正则引擎对跨行预查支持有限,所以第一个版本的兼容性会更好哦。
内容的提问来源于stack exchange,提问作者Jayjay Jay
相关产品推荐
相关产品推荐

