如何阻止Boost的escaped_list_separator消耗带引号令牌的引号?求替代方案
问题解决:保留Boost分词中引号内的嵌套引号
Boost的escaped_list_separator本身就是设计成自动移除引号字符的,而且因为你的语法不支持转义内部引号,它会把内部的{/}误判为引号的开始/结束,导致所有引号都被剥离,没法通过配置修改这个行为。你可以通过以下两种方案实现需求:
方案一:自定义Boost Tokenizer分隔符
自己实现一个符合TokenizerFunc接口的分隔符类,手动控制引号处理逻辑——只移除最外层的{},保留内部的引号和分隔符:
#include <boost/tokenizer.hpp> #include <string> #include <iostream> struct CustomSeparator { using char_type = char; template<typename Iterator> bool operator()(Iterator& next, Iterator end, std::string& token) { token.clear(); if (next == end) return false; // 跳过前置的分号 while (next != end && *next == ';') { ++next; } if (next == end) return false; while (next != end) { if (*next == ';') { ++next; return true; } else if (*next == '{') { // 跳过最外层的左引号 ++next; // 读取直到右引号,保留内部所有字符 while (next != end && *next != '}') { token += *next; ++next; } // 跳过最外层的右引号 if (next != end) { ++next; } } else { token += *next; ++next; } } return true; } // 满足tokenizer接口要求的冗余函数 bool is_escape(char) const { return false; } bool is_quote(char) const { return false; } }; int main() { std::string input("ID=abcde;PARAM={this;{is};quoted}"); boost::tokenizer<CustomSeparator> tokenizer(input); for(const auto &token : tokenizer) { std::cout << token << std::endl; } return 0; }
运行后输出:
ID=abcde PARAM=this;{is};quoted
方案二:用Boost Spirit解析
如果后续语法可能更复杂,用Boost Spirit Qi编写解析规则会更灵活,可读性也更好:
#include <boost/spirit/include/qi.hpp> #include <string> #include <vector> #include <iostream> namespace qi = boost::spirit::qi; int main() { std::string input = "ID=abcde;PARAM={this;{is};quoted}"; std::vector<std::string> tokens; // 定义解析规则:每个token包含普通字符和嵌套的{...},用分号分隔 bool success = qi::parse(input.begin(), input.end(), (qi::char_ - ';' - '{') >> *(('{' >> (qi::char_ - '}') >> '}') | (qi::char_ - ';')) % ';' , tokens); if (success) { for (const auto& token : tokens) { std::cout << token << std::endl; } } else { std::cout << "解析失败" << std::endl; } return 0; }
该代码输出和方案一完全一致,且规则可轻松扩展以支持更多嵌套结构或特殊字符。
内容的提问来源于stack exchange,提问作者sigy
相关产品推荐
相关产品推荐

