如何在Bison中处理未定义、意外token错误并输出token值及行号?
可行性说明
该需求完全可以实现,只需要配合Bison的自定义错误处理配置、Flex的位置跟踪能力、token值跨模块传递三个部分的改造即可完成,最终可输出未识别token的具体内容、对应行号、列号等信息。
具体实现方案
1. 开启Bison错误增强与位置跟踪配置
在你的Bison源码(.y后缀文件)的声明段加入以下配置:
// 输出详细的语法错误信息,替代默认的 generic syntax error %define parse.error verbose // 启用Bison内置的位置跟踪结构体,默认包含行号、列号字段 %locations // 传递自定义上下文参数到解析器和词法分析器,可根据自己的业务扩展 %parse-param {int *custom_lineno} %lex-param {int *custom_lineno} // 声明yylval联合类型,用于传递token文本内容 %union { char* str; } // 声明未定义token的类型 %token <str> UNDEFINED_TOKEN
2. 改造Flex侧代码传递token与位置信息
在你的Flex源码(.l后缀文件)中新增位置更新逻辑,以及未识别字符的兜底处理规则:
%{ #include "y.tab.h" int yylineno = 1; int yycolumn = 1; // 每次匹配token时自动更新位置信息 #define YY_USER_ACTION yylloc.first_line = yylloc.last_line = yylineno; \ yylloc.first_column = yycolumn; yylloc.last_column = yycolumn + yyleng - 1; \ yycolumn += yyleng; %} // 启用Flex内置行号跟踪 %option yylineno %% // 原有你的词法规则保留,在最后新增兜底规则处理未识别字符 . { yylval.str = yytext; // 将未识别的token文本传递到Bison侧 return UNDEFINED_TOKEN; } // 换行时更新行号、重置列号 \n { yylineno++; yycolumn = 1; } %%
3. 自定义yyerror函数处理两类错误
扩展yyerror函数的入参,接收位置信息、token值,区分处理undefined和unexpected token两类错误:
#include <string.h> #include <stdio.h> #include "y.tab.h" void yyerror(YYLTYPE *locp, int *custom_lineno, const char *msg) { // 处理未定义token错误 if (strstr(msg, "$undefined") != NULL || strstr(msg, "undefined") != NULL) { fprintf(stderr, "错误行 %d,列 %d:未识别的token内容为「%s」\n", locp->first_line, locp->first_column, yylval.str); } // 处理非预期token错误 else if (strstr(msg, "unexpected") != NULL) { fprintf(stderr, "错误行 %d,列 %d:语法错误,%s\n", locp->first_line, locp->first_column, msg); } // 其他通用错误 else { fprintf(stderr, "错误行 %d:%s\n", locp->first_line, msg); } }
4. 可选增强配置
如果你使用的是Bison 3.0及以上版本,可以在Bison声明段额外添加%define parse.lac full配置,开启全量Lookahead Correction能力,可输出更完整的错误上下文,与你现有日志中的LAC模块配合获得更精准的错误诊断结果。
效果示例
改造完成后触发错误时会输出类似以下格式的信息:
错误行 15,列 7:未识别的token内容为「#$%」 错误行 22,列 3:语法错误,unexpected EXECSQL, expecting SEMICOLON
内容的提问来源于stack exchange,提问作者Anton Golovenko
相关产品推荐
相关产品推荐

