You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++11正则表达式输出异常:为何跳过首个匹配项?

问题:C++正则分割字符串时跳过首个匹配项并抛出异常的原因

原始代码

#include <iostream>
#include <regex>
#include <string>
#include <vector>
using namespace std;
#define debug(exp) do { cout << #exp << ": " << (exp) << endl; } while (0)
int main (int argc, char *argv[])
{
  vector<string> res;
  string pattern (""),
    text ("ilovechina");
  vector<string> dict = { "i", "love", "china", "lovechina", "ilove" };
  vector<string>::iterator dict_it = dict.begin ();
  if (dict_it == dict.end ()) throw invalid_argument (__func__);
  while (true)
    {
      pattern += *dict_it;
      if (++dict_it == dict.end ()) break;
      else pattern += "|";
    }
  debug (pattern);
  regex re (pattern);
  sregex_iterator ri (text.begin (), text.end (), re),
    ri_end;
  for (; ri != ri_end; ++ri)
    {
      ssub_match sm = (*ri)[0];

      // Edit:
      // Problem here. match_results::position shouldn't
      // be used to check matches.
      // This `if` is originally to check instant matches. 
      // The standard says it's the distance from the target sequence
      // but I misunderstood which the target is.
      // "The string being searched" is the correct name.
      if (ri->position (0)) // The near string does not correspond to the dictionary
    {
      debug (ri->position (0));
      debug (ri->length (0));
      debug (string (sm.first, sm.second));
      throw runtime_error ("not in dictionary");
    }
      // Edit:
      // But I used match_results::position correctly here...
      // This is originally used to check if all words
      // are "instantly" matched.
      if (ri->position (0) + ri->length (0) == text.length ()) // successfully delimited all the words in the text
    break;
      res.push_back (sm.str ());
    }
  if (ri != ri_end) throw runtime_error ("not in dictionary");
}

运行输出

pattern: i|love|china|lovechina|ilove
ri->position (0): 1
ri->length (0): 4
string (sm.first, sm.second): love
terminate called after throwing an instance of 'std::runtime_error'
  what():  not in dictionary
Aborted (core dumped)

问题原因

你的代码存在两个核心错误:

  1. 对match_results::position()的含义理解错误
    position(0)返回的是匹配结果在整个被搜索字符串中的起始偏移量,而非相对于上一次匹配结束位置的偏移。你想用它判断当前匹配是否从剩余字符串的开头开始,但第一次匹配后,后续匹配的偏移量必然大于0,直接触发异常判断。

  2. 循环逻辑错误
    regex_iterator的工作逻辑是:找到当前位置的第一个匹配后,从匹配结束的位置继续查找下一个。你期望每次匹配都从整个字符串的起始位置(偏移0)开始,这完全不符合迭代器的工作机制。

具体到你的测试场景:

  • 第一次匹配到字符串开头的i(偏移0,长度1),代码将其加入res,迭代器移动到偏移1的位置。
  • 第二次匹配到偏移1的love,此时position(0)返回1,触发if (ri->position(0))判断,抛出异常,这就是你看到的输出。

另外,正则分支的匹配顺序也影响结果:你的pattern是i|love|china|lovechina|ilove,ECMAScript模式下的正则引擎会从左到右尝试分支,先匹配到i就停止,不会继续尝试更长的ilove,所以第一个匹配是i而非ilove。


修复方案

修改代码逻辑,跟踪上一次匹配结束的位置,确保每次匹配都从该位置开始,避免跳过字符:

#include <iostream>
#include <regex>
#include <string>
#include <vector>
#include <algorithm>
using namespace std;
#define debug(exp) do { cout << #exp << ": " << (exp) << endl; } while (0)

int main (int argc, char *argv[])
{
    vector<string> res;
    string pattern(""), text("ilovechina");
    vector<string> dict = { "i", "love", "china", "lovechina", "ilove" };
    
    if (dict.empty()) throw invalid_argument(__func__);
    
    // 优化:按字符串长度降序排序,优先匹配长词,避免短词截断长词
    sort(dict.begin(), dict.end(), [](const string& a, const string& b) {
        return a.size() > b.size();
    });
    
    // 构造正则pattern
    for (size_t i = 0; i < dict.size(); ++i) {
        if (i > 0) pattern += "|";
        pattern += dict[i];
    }
    
    debug(pattern);
    regex re(pattern);
    sregex_iterator ri(text.begin(), text.end(), re), ri_end;
    
    size_t last_pos = 0; // 跟踪上一次匹配结束的位置
    for (; ri != ri_end; ++ri) {
        const auto& match = *ri;
        ssub_match sm = match[0];
        
        // 检查当前匹配是否从上次结束的位置开始,防止跳过字符
        if (match.position(0) != last_pos) {
            debug(match.position(0));
            debug(match.length(0));
            debug(string(sm.first, sm.second));
            throw runtime_error("not in dictionary");
        }
        
        res.push_back(sm.str());
        last_pos = match.position(0) + match.length(0);
        
        // 匹配到字符串末尾时提前退出
        if (last_pos == text.length()) {
            break;
        }
    }
    
    // 检查是否完全匹配整个字符串
    if (last_pos != text.length()) {
        throw runtime_error("not in dictionary");
    }
    
    // 输出分割结果
    cout << "分割结果: ";
    for (const auto& s : res) {
        cout << s << " ";
    }
    cout << endl;
    
    return 0;
}

额外优化说明

  • 优先匹配长词:对字典按长度降序排序,避免正则引擎优先匹配短词(比如先匹配ilove而非i),保证分割结果符合预期。
  • 正则转义:如果字典中包含正则特殊字符(如.、*等),需要对这些字符进行转义,否则会导致匹配错误。

内容的提问来源于stack exchange,提问作者dizzy x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 15:42:03