C++正则提取分组内容:regex_iterator问题及转义疑问
提取字符串中的目标内容及正则相关问题解答
一、用regex_iterator提取分组内容
你现在用regex_iterator时输出的是整个匹配(比如{{Test}}),是因为调用了match.str()——这个方法默认返回整个匹配的字符串。而你要的是正则里括号捕获的分组内容,应该用match.str(1),这里的1对应第一个捕获组(也就是正则中(\w+)的部分)。
修改后的代码如下:
#include <regex> #include <iostream> int main() { const std::string s = "<abc>{{Test}}</abc><def>{{Again}}</def>"; std::regex rgx("\\{\\{(\\w+)\\}\\}"); std::sregex_iterator next(s.begin(), s.end(), rgx); std::sregex_iterator end; while (next != end) { std::smatch match = *next; std::cout << match.str(1) << "\n"; // 用str(1)获取第一个捕获组内容 next++; } return 0; }
运行后就能输出Test和Again了。
二、用regex_search匹配多个结果
regex_search默认只会找到第一个匹配项,要匹配所有结果的话,需要循环调用它,并且每次从上一次匹配的结束位置开始继续搜索。具体做法是用match.suffix().first作为下一次搜索的起始位置:
#include <regex> #include <iostream> int main() { const std::string s = "<abc>{{Test}}</abc><def>{{Again}}</def>"; std::regex rgx("\\{\\{(\\w+)\\}\\}"); std::smatch match; auto search_start = s.cbegin(); while (std::regex_search(search_start, s.cend(), match, rgx)) { std::cout << match.str(1) << "\n"; // 取第一个捕获组内容 search_start = match.suffix().first; // 更新搜索起始位置 } }
这样就能输出所有匹配的目标内容了。
三、为什么要用两个反斜杠转义{或}?
这是两层转义叠加的结果:
- C++字符串层面:在C++的字符串字面量中,
\是转义字符(比如\n表示换行、\"表示双引号),所以如果想在字符串里表示一个真正的\,必须写成\\。 - 正则表达式层面:在正则语法中,
{和}是特殊字符,用来表示重复次数(比如a{2,3}表示匹配2到3个a)。如果想匹配字面意义上的{或},就需要用\来转义它们,也就是\{和\}。
把两层转义结合起来,在C++字符串里要写出正则的\{,就需要把\转义成\\,最终变成\\{,同理\\}。
内容的提问来源于stack exchange,提问作者PapaDiHatti
相关产品推荐
相关产品推荐

