如何从C++字符串中提取指定起始词到结束词间的内容?
提取字符串中
begin到end之间的内容 我需要从字符串里提取begin和end之间的内容,示例如下:
string s = "some text \n begin \n text, text, text \n end , some other text"; // 期望输出:\n text, text, text \n
尝试过的无效方法
无效的正则表达式
我写的正则表达式无法正常工作:
std::regex rg(" ^begin: .*[\\S\\s] * ? end.*");
之前可用的正则表达式
之前有一个能捕获从begin:到---整行内容的正则:
std::regex rg(" ^begin: .*[\\S\\s] * ? -{3}.*");
std::copy_if的尝试
我还试过用std::copy_if,但不知道lambda表达式里该写什么逻辑:
std::string result = {}; std::istringstream stream(stringPassed);//用于逐词遍历 std::copy_if(std::istream_iterator<std::string>{stream}, std::istream_iterator<std::string>{}, back_inserter(result), /*此处应编写什么?*/);
解决方案
修正后的正则表达式
原正则存在几个问题:多余的冒号(示例中是begin而非begin:)、空格处理不当、贪婪匹配逻辑错误。可以使用以下带捕获组的正则来提取目标内容:
std::regex rg(R"(begin\s*(.*?)\s*end)", std::regex::dotall); std::smatch match; if (std::regex_search(s, match, rg)) { std::string result = match[1]; // result即为期望提取的内容 }
说明:
R"(...)"是C++11的原始字符串字面量,避免转义字符的繁琐处理std::regex::dotall参数让.匹配包括换行在内的所有字符.*?是非贪婪匹配,确保只匹配到第一个end为止
手动查找位置截取
如果不想使用正则,也可以直接查找begin和end的位置来截取内容:
size_t begin_pos = s.find("begin"); if (begin_pos != std::string::npos) { begin_pos += std::string("begin").length(); size_t end_pos = s.find("end", begin_pos); if (end_pos != std::string::npos) { std::string result = s.substr(begin_pos, end_pos - begin_pos); // result包含begin之后到end之前的所有内容,包括换行和空格 } }
关于std::copy_if的说明
std::istream_iterator<std::string>会按空白符分割逐词读取,会丢失原字符串的换行和空格,且很难精准判断哪些词处于begin和end之间,因此这种方法并不适合你的需求,更推荐上面两种方案。
内容的提问来源于stack exchange,提问作者anon
相关产品推荐
相关产品推荐

