能否在Yacc中定义仅最后测试案例可省略分隔符的语法?
问题描述
作为学习lex和yacc的练习,我编写了一个简单的解析器,用于处理单整数输入的测试用例,输入由仅包含##的单行分隔。示例输入如下:
1 ## 5 ## 21 ##
该解析器已能正常工作,但我希望修改语法,使得仅最后一个测试用例可以省略分隔符##。
我的思路是,如果存在一个可用于规则中的“输入结束”魔法标记,我就可以编写类似如下的规则:
| test_case_suite test_case <magic_end_of_input> { add_case($2); }
但遗憾的是,我并未找到此类标记。
能否在yacc中定义这种语法?
现有代码
Lex 输入
%{ #include "y.tab.h" %} %% [0-9]+ { yylval.i = atoi(yytext); return INTEGER; } ^##$ { return TERMINATOR; } %% int yywrap(void) { return 1; }
Yacc 语法
%{ #include <stdio.h> #include <stdlib.h> int yylex(void); void yyerror(char *); struct TestCase { int n; }; struct Node { struct TestCase* test_case; struct Node* next; }; struct Node* root = NULL; void add_case(struct TestCase* test_case) { struct Node* new_root = malloc(sizeof(struct Node)); new_root->test_case = test_case; new_root->next = root; root = new_root; } %} %union { int i; struct TestCase* tcPtr; }; %token <i> INTEGER %token TERMINATOR %type <tcPtr> test_case %% test_case_suite: test_case_suite test_case TERMINATOR { add_case($2); } | test_case_suite test_case { add_case($2); } | ; test_case: INTEGER { struct TestCase* test_case = malloc(sizeof(struct TestCase)); test_case->n = $1; $$ = test_case; } %% void yyerror(char *s) { fprintf(stdout, "%s\n", s); } int main(void) { yyparse(); struct Node* cur = root; while (cur != NULL) { printf("n = %d\n", cur->test_case->n); cur = cur->next; } return 0; }
解决方案
完全可以实现这种语法,不需要所谓的“输入结束魔法标记”。yacc 的 yyparse() 函数会在 lex 返回 0(表示输入结束)时尝试匹配语法的起始规则,我们只需要调整语法规则,确保只有最后一个测试用例可以省略分隔符,中间的测试用例必须跟分隔符。
修改后的 Yacc 语法规则
将原有的 test_case_suite 规则替换为以下内容即可:
%% test_case_suite: %empty // 允许空输入 | test_case { add_case($1); } // 单个测试用例(无分隔符) | test_case_suite TERMINATOR test_case { add_case($3); } // 已有套件 + 分隔符 + 新测试用例 ; test_case: INTEGER { struct TestCase* test_case = malloc(sizeof(struct TestCase)); test_case->n = $1; $$ = test_case; } %%
规则解释
%empty:允许输入为空的情况;test_case:匹配单个没有后续分隔符的测试用例,对应最后一个测试用例的场景;test_case_suite TERMINATOR test_case:要求每新增一个测试用例前必须有分隔符,确保中间的测试用例都带分隔符,只有最后一个可以省略。
验证效果
修改后,以下输入都能正确解析:
- 带所有分隔符的输入:
1 ## 5 ## 21 ##
- 最后一个测试用例省略分隔符的输入:
1 ## 5 21
而如果中间的测试用例省略分隔符(比如1\n5\n##\n21),解析器会报错,符合需求。
内容的提问来源于stack exchange,提问作者merlin2011
相关产品推荐
相关产品推荐

