C++11 regex_search匹配行为疑问:为何匹配最长而非预期子串?
为什么C++ regex_search中
test 1.*abc只匹配最长结果而非第一个abc结尾的片段? 问题核心出在正则表达式的.*上——它是贪婪匹配量词,会尽可能匹配最多的字符,直到找到字符串里最后一个符合条件的abc,而非第一个。
你的输入字符串是:
test 1 2 3 4 5 abc def abc
当使用test 1.*abc时,test 1匹配开头后,.*会从test 1之后的位置开始,一直匹配到最后一个abc的前一个字符,最终整个匹配结果就是从开头到最后一个abc的完整片段:test 1 2 3 4 5 abc def abc。
如果想要匹配到第一个abc结尾的部分(也就是test 1 2 3 4 5 abc),需要把贪婪匹配改成非贪婪匹配,将.*替换成.*?,对应的正则表达式改为test 1.*?abc。
修改后的代码示例:
#include <iostream> #include <regex> #include <string> using namespace std; int main() { std::string inputStr = "test 1 2 3 4 5 abc def abc"; std::string regexPattern = "test 1.*?abc"; // 改用非贪婪匹配 regex regexp(regexPattern, std::regex::grep); smatch m; while(std::regex_search(inputStr, m, regexp, std::regex_constants::match_default)) { std::cout<<"String that matches the pattern: "<< m.str() << std::endl; inputStr = m.suffix(); } return 0; }
运行后输出会是:
String that matches the pattern: test 1 2 3 4 5 abc String that matches the pattern: def abc
补充说明:你代码里的std::regex::grep标志只是让正则遵循POSIX grep的语法规则,不会改变量词的贪婪/非贪婪特性,所以这个问题的本质还是量词的匹配行为。
内容的提问来源于stack exchange,提问作者bobix
相关产品推荐
相关产品推荐

