PLY词法语法分析器Token匹配不全及语法错误问题求助
Hey there, let's work through your PLY parser issues one by one—those quirks with function detection, global HttpGetText calls, and the VOID syntax error are super common when building C-like parsers with PLY, so let's break down fixes and debugging steps:
1. Fix the VOID Reserved Word Syntax Error
First, let's tackle the low-hanging fruit: the VOID token issue.
- Check your lexer definition: Make sure your
VOIDtoken matches the actual case used in your C code (most C code uses lowercasevoid). For example:def t_VOID(t): r'void' # Match lowercase void, not uppercase VOID unless your code uses that return t - Update your grammar rules: Ensure your yacc parser accepts
VOIDas a valid return type. If your currenttyperule only includes things likeINTorCHAR, addVOIDto it:def p_type(p): '''type : INT | CHAR | VOID''' p[0] = p[1]
This will stop the parser from throwing errors when it encounters void function definitions.
2. Capture All Function Definitions (Not Just main() and func3())
The partial function detection usually stems from gaps in your grammar rules or broken state management in the lexer. Here's what to check:
- Cover all function definition formats: Your current rule might only handle functions with no parameters, or miss nested braces in function bodies. Update the function definition rule to support parameter lists and nested blocks:
# Track current function to link HttpGetText calls global current_function functions = [] # Handle functions with or without parameters def p_function_def(p): '''function_def : type ID '(' param_list ')' '{' statement_list '}' | type ID '(' ')' '{' statement_list '}' ''' current_function = {'name': p[2], 'calls': []} functions.append(current_function) # Support comma-separated parameters def p_param_list(p): '''param_list : param | param_list ',' param ''' def p_param(p): '''param : type ID ''' # Ensure statement_list handles nested braces (like if/else blocks inside functions) def p_statement_list(p): '''statement_list : statement | statement_list statement ''' def p_statement(p): '''statement : expression_statement ';' | '{' statement_list '}' ''' # Nested blocks - Fix lexer state transitions: If you're using custom states in your lexer (like tracking inside/outside functions), make sure you switch states correctly. For example, switch to an
IN_FUNCTIONstate when you hit a{after a function definition, and switch back when you hit the closing}.
3. Capture Global Scope HttpGetText Calls
Right now, your parser is probably only looking for HttpGetText inside function bodies. To catch global calls, add a top-level rule for global statements:
# Store global HttpGetText calls global_global_calls = [] # Top-level program rule: global statements + function definitions def p_program(p): '''program : global_statements function_list | function_list ''' # Handle global scope statements (like function calls) def p_global_statements(p): '''global_statements : global_statement | global_statements global_statement ''' def p_global_statement(p): '''global_statement : expression_statement ';' ''' # Check if this is an HttpGetText call and record it if isinstance(p[1], dict) and p[1]['type'] == 'HttpGetText': global_global_calls.append(p[1]) # Update function call rule to flag HttpGetText def p_function_call(p): '''function_call : ID '(' arg_list ')' | ID '(' ')' ''' if p[1] == 'HttpGetText': call_details = {'type': 'HttpGetText', 'args': p[3] if len(p) > 3 else []} # If we're inside a function, link the call to that function if current_function: current_function['calls'].append(call_details) p[0] = call_details
4. Debugging Tips to Avoid Hacky Workarounds (Like Adding Invalid Characters)
Instead of relying on invalid characters to "reset" the parser, use PLY's built-in debugging tools:
- Enable lexer debugging: Add print statements to your token handlers to verify all tokens are being generated correctly:
def t_ID(t): r'[a-zA-Z_][a-zA-Z0-9_]*' # Check for reserved words t.type = reserved.get(t.value, 'ID') print(f"Generated Token: {t.type} -> {t.value}") # Debug output return t - Enable yacc debugging: Initialize your parser with debug mode to see the parse tree and state transitions:
parser = yacc.yacc(debug=True) - Improve error handling: PLY's default error handling can skip large chunks of code. Add a custom error handler to recover gracefully:
def p_error(p): if p: print(f"Syntax error at token: {p.type} (value: {p.value})") # Skip tokens until we hit a safe point (semicolon or brace) to keep parsing parser.errok() while True: tok = parser.token() if not tok or tok.type in ('SEMICOLON', 'LBRACE', 'RBRACE'): break parser.token() else: print("Syntax error at end of input")
By addressing these areas, you should be able to capture all functions, global HttpGetText calls, and eliminate the VOID syntax error without relying on hacky workarounds.
内容的提问来源于stack exchange,提问作者doublesobig

