如何用boost::spirit::x3实现带转义序列的引号字符串解析器
实现支持完整转义序列的Boost Spirit X3引号字符串解析器
我来帮你搞定这个Boost Spirit X3的字符串解析问题——你原来用的x3::lexeme[quote > *(x3::char_ - quote) > quote]确实只能处理纯普通字符串,完全搞不定转义引号、控制字符这些场景。下面是一套完整的实现,不仅能处理\n、\t这类基础转义,还支持十六进制、八进制甚至ANSI终端转义序列。
完整实现代码
先放上所有依赖和核心代码,你可以直接复用:
#include <boost/spirit/home/x3.hpp> #include <boost/spirit/include/classic_lexer.hpp> #include <boost/spirit/include/classic_escape_char.hpp> #include <stdexcept> #include <string> namespace x3 = boost::spirit::x3; namespace classic = boost::spirit::classic; // 自定义错误类型 struct Error : std::runtime_error { using std::runtime_error::runtime_error; }; int main() { // 定义引号常量 constexpr auto quote = x3::char_('"'); // 转义序列处理Lambda:负责把原始字符串里的转义序列替换成实际字符 auto handle_escape_sequences = [&](auto&& context) -> void { std::string& str = x3::_val(context); uint32_t i{}; // 辅助函数:写入解析后的转义字符 static auto replace = [&](const char replacement) -> void { str[i++] = replacement; }; // 用Spirit Classic的lex_escape_ch_p处理所有转义序列 if (!classic::parse(std::begin(str), std::end(str), *classic::lex_escape_ch_p[replace]).full) { throw Error{ "invalid literal" }; } // 截断字符串到实际有效长度 str.resize(i); }; // 最终的引号字符串解析规则 auto quoted_string = x3::lexeme[ quote > *( // 匹配转义引号\",转换成普通引号存入结果 ("\\\"" >> &x3::char_) >> x3::attr(quote) // 匹配所有非引号的普通字符 | ~x3::char_(quote) ) > quote ][handle_escape_sequences]; // 测试示例 std::string input = R"("Hello \"World\"!\n\tANSI: \033[31mRed Text\033[0m")"; std::string result; try { if (x3::parse(input.begin(), input.end(), quoted_string, result)) { std::cout << "Parsed result:\n" << result << std::endl; } else { std::cerr << "Parse failed!" << std::endl; } } catch (const Error& e) { std::cerr << "Error: " << e.what() << std::endl; } return 0; }
关键部分拆解
1. 解析规则设计
x3::lexeme:确保解析过程中跳过无关空白(如果你的语法需要保留字符串内的空白,这个可以保留,因为它只会跳过字符串外部的空白)- 核心匹配逻辑
*( ... | ... ):("\\\"" >> &x3::char_) >> x3::attr(quote):专门匹配转义引号\",通过x3::attr把它转换成普通引号,这样后续的转义处理能正确识别它~x3::char_(quote):匹配所有非引号的普通字符,直接加入结果字符串
2. 转义序列处理
这里巧妙利用了Spirit Classic的lex_escape_ch_p——它原生支持几乎所有常见转义类型:
- 基础控制字符:
\n、\t、\r、\b等 - 八进制转义:
\0到\777格式的序列 - 十六进制转义:
\x00到\xFF格式的序列 - ANSI终端转义:比如
\033[31m这种用于终端着色、光标定位的序列也能正确解析 - 转义处理Lambda里的
replace函数负责把解析后的字符写入结果,最后截断字符串去除无效部分,解析失败就抛出错误
使用说明
你可以直接把quoted_string规则集成到自己的语法里,测试代码里的示例会解析包含转义引号、换行、制表符和红色ANSI文本的字符串,输出格式化后的内容。
内容的提问来源于stack exchange,提问作者user13501676
相关产品推荐
相关产品推荐

