如何修改词法分析器代码以识别完整多数字符串而非单个数字
解决词法分析器无法识别完整数字字符串的问题
原代码的核心问题是逐个字符遍历输入,导致连续数字被拆分成单个token。要解决这个问题,我们需要改为按token扫描,收集连续的数字/字母作为完整token后再进行分类判断。
修改后的代码如下:
# 移除空格的函数 def remove_spaces(word): return word.replace(" ", "") # 定义token类型的字典 lexical_category = { "+": "Token lexical category: ADDITION", "-": "Token lexical category: SUBTRACTION", "/": "Token lexical category: DIVISION", "*": "Token lexical category: MULTIPLICATION", "%": "Token lexical category: MODULE", } input_str = input("请输入要分析的字符串: ") processed_str = remove_spaces(input_str) index = 0 length = len(processed_str) while index < length: current_char = processed_str[index] # 收集连续数字,生成完整CONSTANT token if current_char.isdigit(): token = current_char index += 1 while index < length and processed_str[index].isdigit(): token += processed_str[index] index += 1 print(f"Given token: {token}") print("Token lexical category: CONSTANT") # 收集连续字母,生成完整VARIABLE token elif current_char.isalpha(): token = current_char index += 1 while index < length and processed_str[index].isalpha(): token += processed_str[index] index += 1 print(f"Given token: {token}") print("Token lexical category: VARIABLE") # 处理运算符类单个字符token else: print(f"Given token: {current_char}") print(lexical_category.get(current_char, "Unknown token")) index += 1
关键改动说明:
- 改用while循环+索引跟踪的遍历方式,主动控制扫描范围,能收集连续的同类字符形成完整token。
- 遇到数字时,会持续向后扫描直至非数字字符,拼接所有连续数字作为一个
CONSTANTtoken。 - 同步优化了变量识别:支持多字母组成的变量名(比如
abc会被识别为单个VARIABLEtoken)。 - 对未知字符增加默认提示,避免输出
None。
测试示例:输入111+abc-222,输出结果为:
Given token: 111 Token lexical category: CONSTANT Given token: + Token lexical category: ADDITION Given token: abc Token lexical category: VARIABLE Given token: - Token lexical category: SUBTRACTION Given token: 222 Token lexical category: CONSTANT
内容的提问来源于stack exchange,提问作者JPtheOne
相关产品推荐
相关产品推荐

