实验室数据公式解析器开发问询:变量/常量/函数识别实现
Hey there! Let's work through building this formula parser you need. Based on your example and the rules you laid out, here's a practical, step-by-step approach to implement it:
The parser needs three key stages: breaking the input into meaningful chunks, structuring those chunks to respect math precedence, then transforming the structure into your target format.
1. Tokenization (Lexical Analysis)
First, split the raw formula string into discrete "tokens" that represent constants, variables, numbers, operators, parentheses, and functions. You can use regular expressions to match each token type:
- Constants: Match patterns like
Constant_Id_(\d+)— capture the numeric ID to replace with|X|later. - Variables: Match
Model_(\w+)— map the variable name (likeAlpha) to its corresponding{m_id}using a lookup table you'll maintain. - Numbers: Match integers or decimals with
\d+(\.\d+)?. - Operators: Match math symbols with
\^|\+|\-|\*|\/. - Parentheses: Match
\(or\)to preserve order of operations. - Functions: Match function names followed by an opening parenthesis (like
string() — you'll handle arguments later when building the structure.
Pro tip: Strip whitespace first to avoid extra tokens messing up your matches.
2. Build an Abstract Syntax Tree (AST)
Tokens alone don't capture the order of operations. Convert the token list into an AST, where each node represents an operation, value, or function call. For example, your sample formula Constant_Id_2^(5 + Model_Alpha) / (Constant_Id_1 * 2) would become a tree where:
- The root is a division node
- Left child is a power node (base: Constant_Id_2, exponent: addition node)
- Right child is a multiplication node (Constant_Id_1, number 2)
This structure ensures you respect math precedence (parentheses first, then exponents, then multiply/divide, then add/subtract) when transforming to the final format.
3. Traverse & Transform the AST
Walk the AST recursively, converting each node to your target syntax:
- Constant nodes: Replace
Constant_Id_Xwith|X| - Variable nodes: Replace
Model_XXXwith{m_id}using your lookup map - Exponent operations: Convert the
^operator to aPOW(a, b)function call (as seen in your sample) - Functions: Keep the function name and arguments intact, just ensure nested tokens are transformed first (e.g.,
string(Model_Alpha, Constant_Id_3)becomesstring({1}, |3|)) - Preserve parentheses: Wrap sub-expressions in parentheses as needed to maintain the original operation order.
Don't forget these critical scenarios:
- Nested functions: Ensure nested calls (like
POW(Constant_Id_1, SIN(Model_Beta))) are parsed and transformed correctly. - Multi-argument functions: Support functions with comma-separated arguments (your
string(,)rule). - Invalid input: Add error handling for malformed formulas (e.g., unclosed parentheses, unknown variables) to give users clear feedback.
Here's a simplified Python snippet to illustrate tokenization (you'd expand this with AST building and transformation):
import re from enum import Enum class TokenType(Enum): CONSTANT = 1 VARIABLE = 2 NUMBER = 3 OPERATOR = 4 PAREN = 5 FUNCTION = 6 # Predefined variable-to-ID map (adjust based on your app's data) VAR_MAP = {"Alpha": 1, "Beta": 2} def tokenize(formula): tokens = [] formula = formula.replace(" ", "") # Clean whitespace pos = 0 patterns = [ (r"Constant_Id_(\d+)", TokenType.CONSTANT), (r"Model_(\w+)", TokenType.VARIABLE), (r"\d+(\.\d+)?", TokenType.NUMBER), (r"\^|\+|\-|\*|\/", TokenType.OPERATOR), (r"\(|\)", TokenType.PAREN), (r"(\w+)\(", TokenType.FUNCTION) ] while pos < len(formula): matched = False for pattern, token_type in patterns: match = re.match(pattern, formula[pos:]) if match: if token_type == TokenType.CONSTANT: tokens.append((token_type, match.group(1))) elif token_type == TokenType.VARIABLE: var_id = VAR_MAP.get(match.group(1), "unknown") tokens.append((token_type, var_id)) elif token_type == TokenType.FUNCTION: tokens.append((token_type, match.group(1))) else: tokens.append((token_type, match.group(0))) pos += len(match.group(0)) matched = True break if not matched: raise ValueError(f"Invalid character at position {pos}: {formula[pos]}") return tokens # Example usage: sample_formula = "Constant_Id_2^(5 + Model_Alpha) / (Constant_Id_1 * 2)" tokens = tokenize(sample_formula) print(tokens)
If you don't want to build everything from scratch, use a parsing library to handle the heavy lifting:
- Python:
lark-parserorpyparsinglet you define formal grammar rules and auto-generate ASTs. - Java/C#:
ANTLRis a powerful parser generator for complex syntax. - JavaScript:
pegjsorchevrotainwork well for browser/Node.js environments.
These tools reduce boilerplate and ensure you handle edge cases like nested expressions and operator precedence correctly.
内容的提问来源于stack exchange,提问作者Christian4423

