C++11正则表达式输出异常:为何跳过首个匹配项?
问题:C++正则分割字符串时跳过首个匹配项并抛出异常的原因
原始代码
#include <iostream> #include <regex> #include <string> #include <vector> using namespace std; #define debug(exp) do { cout << #exp << ": " << (exp) << endl; } while (0) int main (int argc, char *argv[]) { vector<string> res; string pattern (""), text ("ilovechina"); vector<string> dict = { "i", "love", "china", "lovechina", "ilove" }; vector<string>::iterator dict_it = dict.begin (); if (dict_it == dict.end ()) throw invalid_argument (__func__); while (true) { pattern += *dict_it; if (++dict_it == dict.end ()) break; else pattern += "|"; } debug (pattern); regex re (pattern); sregex_iterator ri (text.begin (), text.end (), re), ri_end; for (; ri != ri_end; ++ri) { ssub_match sm = (*ri)[0]; // Edit: // Problem here. match_results::position shouldn't // be used to check matches. // This `if` is originally to check instant matches. // The standard says it's the distance from the target sequence // but I misunderstood which the target is. // "The string being searched" is the correct name. if (ri->position (0)) // The near string does not correspond to the dictionary { debug (ri->position (0)); debug (ri->length (0)); debug (string (sm.first, sm.second)); throw runtime_error ("not in dictionary"); } // Edit: // But I used match_results::position correctly here... // This is originally used to check if all words // are "instantly" matched. if (ri->position (0) + ri->length (0) == text.length ()) // successfully delimited all the words in the text break; res.push_back (sm.str ()); } if (ri != ri_end) throw runtime_error ("not in dictionary"); }
运行输出
pattern: i|love|china|lovechina|ilove ri->position (0): 1 ri->length (0): 4 string (sm.first, sm.second): love terminate called after throwing an instance of 'std::runtime_error' what(): not in dictionary Aborted (core dumped)
问题原因
你的代码存在两个核心错误:
对
match_results::position()的含义理解错误position(0)返回的是匹配结果在整个被搜索字符串中的起始偏移量,而非相对于上一次匹配结束位置的偏移。你想用它判断当前匹配是否从剩余字符串的开头开始,但第一次匹配后,后续匹配的偏移量必然大于0,直接触发异常判断。循环逻辑错误
regex_iterator的工作逻辑是:找到当前位置的第一个匹配后,从匹配结束的位置继续查找下一个。你期望每次匹配都从整个字符串的起始位置(偏移0)开始,这完全不符合迭代器的工作机制。
具体到你的测试场景:
- 第一次匹配到字符串开头的
i(偏移0,长度1),代码将其加入res,迭代器移动到偏移1的位置。 - 第二次匹配到偏移1的
love,此时position(0)返回1,触发if (ri->position(0))判断,抛出异常,这就是你看到的输出。
另外,正则分支的匹配顺序也影响结果:你的pattern是i|love|china|lovechina|ilove,ECMAScript模式下的正则引擎会从左到右尝试分支,先匹配到i就停止,不会继续尝试更长的ilove,所以第一个匹配是i而非ilove。
修复方案
修改代码逻辑,跟踪上一次匹配结束的位置,确保每次匹配都从该位置开始,避免跳过字符:
#include <iostream> #include <regex> #include <string> #include <vector> #include <algorithm> using namespace std; #define debug(exp) do { cout << #exp << ": " << (exp) << endl; } while (0) int main (int argc, char *argv[]) { vector<string> res; string pattern(""), text("ilovechina"); vector<string> dict = { "i", "love", "china", "lovechina", "ilove" }; if (dict.empty()) throw invalid_argument(__func__); // 优化:按字符串长度降序排序,优先匹配长词,避免短词截断长词 sort(dict.begin(), dict.end(), [](const string& a, const string& b) { return a.size() > b.size(); }); // 构造正则pattern for (size_t i = 0; i < dict.size(); ++i) { if (i > 0) pattern += "|"; pattern += dict[i]; } debug(pattern); regex re(pattern); sregex_iterator ri(text.begin(), text.end(), re), ri_end; size_t last_pos = 0; // 跟踪上一次匹配结束的位置 for (; ri != ri_end; ++ri) { const auto& match = *ri; ssub_match sm = match[0]; // 检查当前匹配是否从上次结束的位置开始,防止跳过字符 if (match.position(0) != last_pos) { debug(match.position(0)); debug(match.length(0)); debug(string(sm.first, sm.second)); throw runtime_error("not in dictionary"); } res.push_back(sm.str()); last_pos = match.position(0) + match.length(0); // 匹配到字符串末尾时提前退出 if (last_pos == text.length()) { break; } } // 检查是否完全匹配整个字符串 if (last_pos != text.length()) { throw runtime_error("not in dictionary"); } // 输出分割结果 cout << "分割结果: "; for (const auto& s : res) { cout << s << " "; } cout << endl; return 0; }
额外优化说明
- 优先匹配长词:对字典按长度降序排序,避免正则引擎优先匹配短词(比如先匹配
ilove而非i),保证分割结果符合预期。 - 正则转义:如果字典中包含正则特殊字符(如
.、*等),需要对这些字符进行转义,否则会导致匹配错误。
内容的提问来源于stack exchange,提问作者dizzy x
相关产品推荐
相关产品推荐

