You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

regex_token_iterator<>偶现匹配子串丢失问题排查求助

关于regex_token_iterator<>丢失匹配子串的问题分析

问题描述

使用regex_token_iterator<>提取行内所有匹配子串时,代码偶尔会丢失第二个匹配项,且出错的行在不同运行时随机变化。编译器为Apple clang version 14.0.0 (clang-1400.0.29.202),编译参数-std=c++14。改用while循环重复调用regex_search()的版本运行正常,需明确regex_token_iterator<>版本失效的原因。

原代码

#include<regex>
#include<iostream>
#include<string>
#include<fstream>
#include<sstream>

using namespace std;

struct bad_from_string : bad_cast{
  const char* what() const noexcept override{
    return "bad cast from string";
  }
};

template<typename T>
T from_string(const string& s){
  istringstream is{s};
  T t;
  if(!(is>>t))
    throw bad_from_string{};
  return t;
}

int main(){
  regex pat{R"((\d{1,2})/(\d{1,2})/(\d{4}))"}; // e.g. 7/21/2022
  ifstream ifs{"test_regex_token_iterator.txt"};
  ofstream ofs{"test_out_regex_token_iterator.txt"};

  regex_token_iterator<string::iterator> rend; // default constructor is used for indicating the end of the sequence
  
  for(string line; getline(ifs, line);){
    smatch matches;
    
    string replace_pattern; 

    int month{0}, day{0}, year{0};

    regex_token_iterator<string::iterator> riter(line.begin(), line.end(), pat);
      
    // for each matched substring, replace it individually
    while(riter!=rend){
      string matched_substring{(*riter).str()};
      // *riter returns a reference to the sub_match object riter is pointing to.
      // sub_match is not a string. sub_match::str() returns the string of the sub_match.
      
      // put each matched substring into variable "matches"
      regex_search(matched_substring, matches, pat);
      
      // get the day, month, and year values in int
      day = from_string<int>(matches.str(2));
      month = from_string<int>(matches.str(1));
      year = from_string<int>(matches.str(3));
      
      // here make replace_pattern yyyy-mm-dd
      if(month<10 && day<10)
        replace_pattern = to_string(year)+"-0"+to_string(month)+"-0"+to_string(day); // both day and month need the fron '0'
      else if(month<10)
        replace_pattern = to_string(year)+"-0"+to_string(month)+"-"+to_string(day);
      else if(day<10)
        replace_pattern = to_string(year)+"-"+to_string(month)+"-0"+to_string(day);
      else
        replace_pattern = to_string(year)+"-"+to_string(month)+"-"+to_string(day);
      
      line = regex_replace(line, regex(matched_substring), replace_pattern); // regex_replace() returns a string
      // since I want to replace only 1 matched substring *riter, I use the exact substring 
      // in the place of regex pattern
      
      ++riter;      // move to the next matched substring
    }
  
    ofs << line << endl; 
  }
  
  return 0;
}

测试文件内容

12/01/2022 - 12/31/2022
12/01/2022 - 12/31/2022
12/01/2022 - 12/31/2022
12/01/2022 - 12/31/2022

10/01/2022 - 10/31/2022
10/01/2022 - 10/31/2022
10/01/2022 - 10/31/2022
10/01/2022 - 10/31/2022
10/01/2022 - 10/31/2022

异常输出(随机出现)

2022-12-01 - 12/31/2022
2022-12-01 - 2022-12-31
2022-12-01 - 12/31/2022
2022-12-01 - 12/31/2022

2022-10-01 - 10/31/2022
2022-10-01 - 2022-10-31
2022-10-01 - 10/31/2022
2022-10-01 - 10/31/2022
2022-10-01 - 10/31/2022

预期输出

2022-12-01 - 2022-12-31
2022-12-01 - 2022-12-31
2022-12-01 - 2022-12-31
2022-12-01 - 2022-12-31

2022-10-01 - 2022-10-31
2022-10-01 - 2022-10-31
2022-10-01 - 2022-10-31
2022-10-01 - 2022-10-31
2022-10-01 - 2022-10-31

原因分析

这不是regex_token_iterator<>的Bug,完全是代码逻辑错误,核心问题有两点:

  1. 迭代器依赖的原始字符串被修改,导致迭代器失效
    regex_token_iterator<>是基于初始的line字符串创建的,它内部保存了指向原字符串内存的迭代器位置。当执行line = regex_replace(...)时,原字符串被替换为新的字符串(内存地址和内容都发生变化),但riter仍然指向旧字符串的内存区域,后续的++riter操作属于未定义行为,因此会随机丢失匹配项。

  2. 替换逻辑不符合预期
    使用regex(matched_substring)作为替换的正则表达式,存在两个问题:

    • 如果匹配子串包含正则特殊字符(虽然本例中日期没有,但场景扩展后会出错),会导致正则解析异常;
    • regex_replace默认会替换所有匹配项,而非当前迭代器指向的单个匹配,这会导致逻辑偏差。

另外,代码中对matched_substring再次调用regex_search属于冗余操作,regex_token_iterator本身可以直接获取捕获组信息。

修正方案

核心原则是不要在遍历迭代器期间修改迭代器依赖的原始字符串,可以先收集所有匹配信息,再统一替换。以下是修正后的代码:

#include<regex>
#include<iostream>
#include<string>
#include<fstream>
#include<sstream>
#include<vector>
#include<algorithm>

using namespace std;

struct bad_from_string : bad_cast{
  const char* what() const noexcept override{
    return "bad cast from string";
  }
};

template<typename T>
T from_string(const string& s){
  istringstream is{s};
  T t;
  if(!(is>>t))
    throw bad_from_string{};
  return t;
}

int main(){
  regex pat{R"((\d{1,2})/(\d{1,2})/(\d{4}))"};
  ifstream ifs{"test_regex_token_iterator.txt"};
  ofstream ofs{"test_out_regex_token_iterator.txt"};

  regex_iterator<string::iterator> rend;

  for(string line; getline(ifs, line);){
    string new_line = line;
    // 收集所有需要替换的位置和内容
    vector<pair<size_t, pair<string, string>>> replace_list;

    regex_iterator<string::iterator> riter(line.begin(), line.end(), pat);
    while(riter != rend){
      const smatch& match = *riter;
      int month = from_string<int>(match[1].str());
      int day = from_string<int>(match[2].str());
      int year = from_string<int>(match[3].str());

      // 构造替换字符串
      string replace_pattern;
      if(month < 10 && day < 10)
        replace_pattern = to_string(year) + "-0" + to_string(month) + "-0" + to_string(day);
      else if(month < 10)
        replace_pattern = to_string(year) + "-0" + to_string(month) + "-" + to_string(day);
      else if(day < 10)
        replace_pattern = to_string(year) + "-" + to_string(month) + "-0" + to_string(day);
      else
        replace_pattern = to_string(year) + "-" + to_string(month) + "-" + to_string(day);

      // 记录原字符串中的匹配位置、原字符串、替换字符串
      replace_list.emplace_back(match.position(), make_pair(match.str(), replace_pattern));
      ++riter;
    }

    // 从后往前替换,避免前面的替换影响后续匹配的位置
    reverse(replace_list.begin(), replace_list.end());
    for(const auto& item : replace_list){
      size_t pos = item.first;
      const string& old_str = item.second.first;
      const string& new_str = item.second.second;
      new_line.replace(pos, old_str.size(), new_str);
    }

    ofs << new_line << endl;
  }

  return 0;
}

关键改进点

  • 使用regex_iterator替代regex_token_iterator,直接获取完整的smatch信息,避免冗余的regex_search调用;
  • 复制原字符串到new_line,遍历原字符串收集所有匹配信息,不修改原字符串,避免迭代器失效;
  • 从后往前执行替换操作,防止前面的替换改变字符串长度,影响后续匹配的位置偏移。

内容的提问来源于stack exchange,提问作者taka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 19:31:03