You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用C++ regex库分词含单词的数学表达式无输出,求解决方法

Fixing Tokenizing of Math Expressions with Words in C++ Regex

Got it, let’s work through this tokenizing problem together. I’ve struggled with similar regex-based tokenization tasks before, so here’s how to get it right.

First, the Most Common Pitfall

The biggest mistake here is wrong regex pattern order. If you match single characters (like +, -, or letters) before matching full words or numbers, your regex will split words into individual letters instead of capturing them as whole tokens. For example, if your pattern first matches [a-z], the word add would turn into a, d, d instead of one add token.

The Right Regex Pattern

You need to prioritize longer, more specific tokens first:

  1. Words/Identifiers: Match sequences of letters, numbers, and underscores (e.g., function names like add, variables like x).
  2. Numbers: Match both integers (e.g., 10) and floating-point numbers (e.g., 5.2).
  3. Operators/Delimiters: Match single-character tokens like +, -, *, /, (, ), ,, =.

Here’s a robust pattern using C++ raw string literals (to avoid messy escape characters):

R"((\w+)|(\d+\.\d+|\d+)|([+\-*/(),=]))"
  • (\w+): Captures words/identifiers (matches [a-zA-Z0-9_]+)
  • (\d+\.\d+|\d+): Captures floats first, then integers
  • ([+\-*/(),=]): Captures all single-character math operators and delimiters

Full C++ Implementation Example

Here’s a complete function that tokenizes your expression correctly, plus a test case:

#include <iostream>
#include <regex>
#include <vector>
#include <string>

using namespace std;

vector<string> tokenize_math_expr(const string& expr) {
    // Regex pattern: prioritize words > numbers > operators/delimiters
    regex token_pattern(R"((\w+)|(\d+\.\d+|\d+)|([+\-*/(),=]))");
    vector<string> tokens;

    // Iterate over all matches in the expression
    sregex_iterator it(expr.begin(), expr.end(), token_pattern);
    sregex_iterator end;

    for (; it != end; ++it) {
        // Grab the first non-empty capture group (alternations are checked in order)
        if (!(*it)[1].str().empty()) {
            tokens.push_back((*it)[1].str());
        } else if (!(*it)[2].str().empty()) {
            tokens.push_back((*it)[2].str());
        } else if (!(*it)[3].str().empty()) {
            tokens.push_back((*it)[3].str());
        }
    }

    return tokens;
}

int main() {
    // Test with a sample expression
    string expr = "3 * add(5.2, x) - subtract(y, 10)";
    vector<string> tokens = tokenize_math_expr(expr);

    cout << "Tokenized Result:\n";
    for (const auto& token : tokens) {
        cout << "\"" << token << "\" ";
    }
    cout << endl;

    return 0;
}

Expected Output

Tokenized Result:
"3" "*" "add" "(" "5.2" "," "x" ")" "-" "subtract" "(" "y" "," "10" ")" 

Debugging Tips If It Still Fails

  • Print intermediate matches: Add a line inside the loop to print each capture group’s value—this will show you which part of the regex is matching (or not matching) your tokens.
  • Handle edge cases: If you need to support negative numbers (e.g., -5.2), adjust the number capture group to (-?\d+\.\d+|-?\d+). Just note that this will capture unary minus signs as part of the number; if you need to distinguish unary minus from subtraction, you’ll need extra logic after tokenization.
  • Check regex flags: If your expression has uppercase letters, add regex::icase to the regex constructor to make matching case-insensitive.

内容的提问来源于stack exchange,提问作者Soner from The Ottoman Empire

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:18:05