You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从C++字符串中提取指定起始词到结束词间的内容?

提取字符串中begin到end之间的内容

我需要从字符串里提取begin和end之间的内容,示例如下:

string s = "some text \n begin \n text, text, text \n end , some other text";
// 期望输出:\n text, text, text \n

尝试过的无效方法

无效的正则表达式

我写的正则表达式无法正常工作:

std::regex rg(" ^begin: .*[\\S\\s] * ? end.*");

之前可用的正则表达式

之前有一个能捕获从begin:到---整行内容的正则:

std::regex rg(" ^begin: .*[\\S\\s] * ? -{3}.*");

std::copy_if的尝试

我还试过用std::copy_if,但不知道lambda表达式里该写什么逻辑:

std::string result = {};
std::istringstream stream(stringPassed);//用于逐词遍历
std::copy_if(std::istream_iterator<std::string>{stream}, 
std::istream_iterator<std::string>{}, back_inserter(result), /*此处应编写什么?*/);

解决方案

修正后的正则表达式

原正则存在几个问题:多余的冒号(示例中是begin而非begin:)、空格处理不当、贪婪匹配逻辑错误。可以使用以下带捕获组的正则来提取目标内容:

std::regex rg(R"(begin\s*(.*?)\s*end)", std::regex::dotall);
std::smatch match;
if (std::regex_search(s, match, rg)) {
    std::string result = match[1];
    // result即为期望提取的内容
}

说明:

  • R"(...)"是C++11的原始字符串字面量,避免转义字符的繁琐处理
  • std::regex::dotall参数让.匹配包括换行在内的所有字符
  • .*?是非贪婪匹配,确保只匹配到第一个end为止

手动查找位置截取

如果不想使用正则,也可以直接查找begin和end的位置来截取内容:

size_t begin_pos = s.find("begin");
if (begin_pos != std::string::npos) {
    begin_pos += std::string("begin").length();
    size_t end_pos = s.find("end", begin_pos);
    if (end_pos != std::string::npos) {
        std::string result = s.substr(begin_pos, end_pos - begin_pos);
        // result包含begin之后到end之前的所有内容,包括换行和空格
    }
}

关于std::copy_if的说明

std::istream_iterator<std::string>会按空白符分割逐词读取,会丢失原字符串的换行和空格,且很难精准判断哪些词处于begin和end之间,因此这种方法并不适合你的需求,更推荐上面两种方案。

内容的提问来源于stack exchange,提问作者anon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 12:45:13