C++如何从DBC格式VAL_字符串中提取正则匹配值至数组
C++ 提取目标参数的解决方案
核心思路
分两步处理:
- 先提取开头的固定参数(数字234、标识符State1),并分离出数组部分的原始内容;
- 对数组内容单独匹配,遍历提取每一组(无符号整数+带转义的字符串)。
代码实现
#include <iostream> #include <string> #include <vector> #include <regex> #include <utility> int main() { // 原始输入字符串(注意转义处理) std::string input = R"(VAL_ 234 State1 123 "Description 1" 0 "Description 2 with \n new line" 90903489 "Big value and special characters &$§())!" ;)"; // 第一步:匹配开头的固定参数和数组内容 std::regex header_re(R"(^VAL_ (\d+) ([A-Za-z_]\w*) (.*?);$)"); std::smatch header_match; if (!std::regex_match(input, header_match, header_re)) { std::cout << "输入格式不符合要求" << std::endl; return 1; } // 提取前两个固定参数 int target_num = std::stoi(header_match[1].str()); std::string target_id = header_match[2].str(); std::string array_raw = header_match[3].str(); // 第二步:遍历提取数组的每一组元素 std::regex item_re(R"((\d+) "((?:\\.|[^"])*)")"); std::vector<std::pair<int, std::string>> array_items; std::sregex_iterator item_it(array_raw.begin(), array_raw.end(), item_re); std::sregex_iterator end_it; for (; item_it != end_it; ++item_it) { std::smatch item_match = *item_it; int item_num = std::stoi(item_match[1].str()); std::string item_desc = item_match[2].str(); array_items.emplace_back(item_num, item_desc); } // 输出验证结果 std::cout << "提取的数字: " << target_num << "\n"; std::cout << "提取的标识符: " << target_id << "\n"; std::cout << "数组内容:\n"; for (const auto& item : array_items) { std::cout << "- " << item.first << " \"" << item.second << "\"\n"; } return 0; }
关键说明
为什么之前的正则只拿到最后一组?
你之前使用的([0-9]*\\s\"[^\"]*\"\\s)+属于重复捕获组,C++标准库的正则引擎只会保留该组最后一次匹配的结果,无法获取所有重复组内容。如何遍历子匹配?
使用std::sregex_iterator可以遍历所有匹配的结果,每个迭代器返回的std::smatch对象包含当前匹配的所有捕获组(比如示例中的组1是整数,组2是描述字符串),通过遍历迭代器就能拿到所有数组元素。处理带转义的字符串
正则((?:\\.|[^"])*)用来匹配带转义字符的字符串:\\.匹配转义字符(比如\n、\");[^"]匹配非引号的普通字符;(?:...)是非捕获组,避免额外的捕获开销。
内容的提问来源于stack exchange,提问作者Murmi
相关产品推荐
相关产品推荐

