Bison语法问题:如何处理语句后可选';'及顶层函数无分号需求
问题与解决方案
需求说明
语句通常以;分隔,但需满足:
- 顶层函数定义(如
fun a() {})无需以;结尾 - 作为表达式一部分的lambda(如
F=fun(){})属于表达式,需跟随;
核心解决方案:语法层面区分语句类型
不需要修改Flex扫描器的状态,直接在Bison语法规则中区分顶层函数定义和表达式语句即可,这是最清晰且无歧义的实现方式:
- 新增
func_def规则作为独立的语句类型,无需;后缀 - 保留原有的
expr ';'作为表达式语句规则 - 函数定义同时作为
primary的一部分,确保能在表达式中使用
修改后的代码文件
Makefile
all: bison -rall -o parse.c n.y flex -o scan.c n.l gcc -g -c parse.c -o parse.o gcc -g -c scan.c -o scan.o gcc -g -c n.c -o n.o gcc -g scan.o parse.o n.o -lc -o n.exe ./n.exe test.txt
n.y(Bison语法文件)
%{ #define YYDEBUG 1 %} %code requires { extern enum yytokentype yylex(); extern void yyerror(const char* errmsg); extern void yyerrorf(const char* format, ...); } %expect 0 %define parse.trace %verbose %header %define parse.error verbose %token FUNC %token ID NUM %left '+' %% %start unit; unit: stmts stmts: stmt {} | stmts stmt {} stmt: expr ';' {} // 表达式语句,需要;结尾 | func_def {} // 顶层函数定义,不需要;结尾 ; func_def: FUNC '(' ')' '{' stmts '}' {} | FUNC ID '(' ')' '{' stmts '}' {} ; expr: expr '+' expr {} | primary {} ; primary: NUM {} | ID {} | func_def {} // 函数定义作为表达式的primary部分 ; %% void yyerror(const char* errmsg) { printf("%s",errmsg); }
n.l(Flex词法文件)
%option noyywrap nounput noinput batch debug %{ #include "parse.h" %} id [a-zA-Z][a-zA-Z_0-9]* int [0-9]+ blank [ \t\r] %% [\n] { } [ \t] { } [#][^\n]* { } fun { return FUNC; } "{" {return '{';} "}" {return '}';} "(" {return '(';} ")" {return ')';} "+" {return '+';} ";" {return ';';} {id} {return ID; } {int} {return NUM; } %%
n.c(主程序文件)
#include <stdlib.h> #include <stdio.h> #include <getopt.h> #include <unistd.h> #include <stdarg.h> #include <ctype.h> #include <sys/stat.h> #include <fcntl.h> #include "parse.h" #define fail(fmt, ...) do { printf(fmt, ##__VA_ARGS__); exit(1); } while(0) static void readin(char *fname) { struct stat st; int fd; char *ptr, *p, *pend; fd = open(fname, O_RDONLY, 0); if (fd < 0) fail("fail open %s", fname); if (fstat(fd, &st) < 0) fail("fail sta %s", fname); if (!S_ISREG(st.st_mode)) { fail("fail %s is not a regular file", fname); } p = ptr = (char *)malloc(st.st_size+1); if (read(fd, ptr, st.st_size) != st.st_size) { } ptr[st.st_size] = 0; extern void * yy_scan_string (const char * yystr ); yy_scan_string( ptr ); yydebug=1; yyparse(); } int main(int argc, char **argv) { if (argc >= 2) { readin(argv[1]); } return 0; }
测试用例
期望通过的测试文件(test.txt)
fun a() { 1; } 1+1;1+fun(){3;};2+2;
运行编译命令后,该测试文件可正常解析:顶层函数块后无需加;,表达式中的lambda仍需跟随;。
关于Flex状态修改的补充说明
你提到的修改Flex在块结尾返回'\n'作为分隔符的思路不可行,原因如下:
- 会引入歧义:换行符在代码中可能只是格式换行,并非语句分隔
- 无法区分顶层和表达式内的函数块,容易导致语法错误
- 语法层面的解决方案更简洁可靠,无需依赖词法分析器的状态切换
内容的提问来源于stack exchange,提问作者Konrad Eisele
相关产品推荐
相关产品推荐

