You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于《Programming: Principles and Practice Using C++》第6章Tokens与Grammar的技术疑问

Understanding Tokens and Grammar for the Calculator in Programming: Principles and Practice Using C++ Chapter 6

Hey there! I totally get where you're coming from—Stroustrup's book can skip over some foundational parsing concepts assuming you might have touched on them before, but let's break this down clearly with the calculator example to make it click.

What Are Tokens, Exactly?

Think of tokens as the "building blocks" of your input. When you type something like 123 + 45*(6-7), the raw string is just a sequence of characters. Tokens are the process of splitting that string into meaningful, indivisible units that your program can work with.

For the calculator, these units would be:

  • Numbers (like 123, 45)
  • Operators (+, -, *, /)
  • Parentheses ((, ))
  • A special "end" token to signal the end of input

So that example string would get split into this list of tokens:
[123, +, 45, *, (, 6, -, 7, ), end]

The point of tokenization is to strip out irrelevant stuff (like spaces or newlines) and group characters into things that have a clear purpose. Stroustrup's code probably has a Token struct or class that holds two things: a type (to say if it's a number, plus sign, etc.) and a value (only used for numbers, to store their actual numeric value). Here's a simplified version of what that might look like:

enum class TokenType { Number, Plus, Minus, Multiply, Divide, LParen, RParen, End };

struct Token {
    TokenType type;
    double value; // Only valid if type is TokenType::Number
};

Your tokenizer function (often called get_token() or similar) would read characters from the input one by one, decide if it's starting a number, an operator, etc., and spit out the next token each time it's called.

Grammar: The Rules for Building Valid Expressions

If tokens are the blocks, grammar is the instruction manual that tells you how to stack those blocks to make a valid expression. It's a set of rules that define what counts as a legal input (and also handles things like operator precedence—why 2+3*4 is 14, not 20).

For the calculator, the grammar might look something like this (written in a simplified form):

  • Expression: Term ( + Term | - Term )*
    (An expression is one or more terms added or subtracted together)
  • Term: Factor ( * Factor | / Factor )*
    (A term is one or more factors multiplied or divided together)
  • Factor: Number | ( Expression )
    (A factor is either a single number, or an expression wrapped in parentheses)

These rules are directly mapped to recursive functions in the code—this is called a recursive descent parser. For example:

  • parse_expression() calls parse_term() repeatedly, checking for + or - tokens between them
  • parse_term() calls parse_factor() repeatedly, checking for * or / tokens
  • parse_factor() handles numbers and parentheses, and calls parse_expression() again if it sees an opening parenthesis

Here's a quick snippet of how that might look:

double parse_expression();
double parse_term();
double parse_factor() {
    Token t = get_token();
    if (t.type == TokenType::Number) {
        return t.value;
    } else if (t.type == TokenType::LParen) {
        double result = parse_expression();
        Token closing_paren = get_token();
        if (closing_paren.type != TokenType::RParen) {
            // Handle error: missing closing parenthesis
        }
        return result;
    }
    // Handle error: unexpected token
}

This recursive structure is exactly how the calculator enforces operator precedence—since parse_expression relies on parse_term, which relies on parse_factor, multiplication and division get evaluated before addition and subtraction automatically.

Why This Matters for Your Project

Tokens let your program stop dealing with raw characters and start dealing with meaningful concepts. Grammar gives you a structured way to interpret those tokens correctly, so you don't end up miscalculating expressions because you didn't account for precedence or parentheses.

Stroustrup might have glossed over some of this because recursive descent parsing and tokenization are standard parsing techniques, but once you connect the token list to the grammar rules, it all starts to make sense.

内容的提问来源于stack exchange,提问作者Constant Furstenberg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:39:07