关于多数C/C++编译器数组词法单元生成的两类技术问询
Answers to Your C/C++ Compiler Lexer Questions
1. Lexical Tokens for MyArray[20] and Array Size Handling
- First off, remember that the lexer (scanner) only has one job: split source code into atomic tokens, without understanding grammar or context. So for
MyArray[20], it will generate four separate, independent tokens:IDENTIFIER(forMyArray)LBRACKET(for[)INTEGER_LITERAL(for20, with its numeric value attached as metadata)RBRACKET(for])
There’s no combined "array_token" orarray_token[const_int]—lexers don’t group tokens based on syntax meaning. That’s the parser’s responsibility later on.
- As for the array size: the lexer doesn’t care what
20is being used for. It just flags it as an integer literal. The actual handling of the array size (like checking if it’s a valid non-negative constant in a declaration, or evaluating it as a runtime value in an array access) happens in the syntax analysis or semantic analysis phases. For example, inint MyArray[20];, the parser will recognize this as an array declaration, then pass the integer literal to the semantic analyzer to verify it meets the language’s rules for array dimensions.
2. Handling MyArray[20.5] in Execution Code
- The lexer still processes this as a sequence of valid tokens:
IDENTIFIER(forMyArray)LBRACKET(for[)FLOATING_LITERAL(for20.5, with its value stored)RBRACKET(for])
The lexer won’t throw an error here because each character sequence is a valid token on its own.
- The problem gets caught later, during semantic analysis. C/C++ requires array subscripts to be an integer type (or a type that can be implicitly converted to integer, like
boolorchar). A floating-point literal like20.5doesn’t fit this requirement, so the compiler will emit an error (typically something like "array subscript is not an integer"). Some compilers might issue a warning first if implicit conversion is allowed, but most will flag this as an error in strict mode.
内容的提问来源于stack exchange,提问作者OneAndOnly
相关产品推荐
相关产品推荐

