C++ Enum值异常:运算符被错误识别为关键字求助
问题描述
我正在为命令行项目开发一款简易解释器,近期遇到了一个无法定位原因的问题。
相关代码如下:
// cls.h #pragma once #include <iostream> #include <string> #include <iomanip> enum id {undef, kwd, idt, opr, val}; // Ignore idt extern std::string keywords[]; // Declaring the keywords array extern std::string operators[]; // Declaring the operators array class Token { id ID; std::string value; public: Token() { ID = undef; value = ""; } void SetID(id ID) { this->ID = ID; } void SetValue(std::string value) { this->value = value; } id GetID() { return ID; } std::string GetValue() { return value; } }; std::string UserInput(); void Lexical(Token *token, int nstr, std::string str, int ntok); // cls.cpp #include <iostream> // Definitions of the keywords and operators arrays std::string keywords[] = { "echo", "exit" }; std::string operators[] = { "+", "-", "/", "%" }; // input.cpp #include "cls.h" std::string UserInput() // Function to take user's input { std::string input; std::getline(std::cin, input); return input; } // lex.cpp #include "cls.h" // To check whether a token is a keyword static bool aKeyword(std::string str) { for (int i = 0; i < 10; i++) { if (keywords[i] == str) return true; } return false; } // To check whether a token is an operator static bool aOperator(std::string str) { for (int i = 0; i < 10; i++) { if (operators[i] == str) return true; } return false; } void Lexical(Token *token, int ntok, std::string str, int nstr) { // Working the tokens' values int idx = 0; bool fspace = false; std::string plh[256]; for (int i = 0; i < nstr; i++) { if ((str.at(i) == ' ') && (fspace == false)) { idx++; fspace = true; continue; } else if ((str.at(i) == ' ') && (fspace == true)) { continue; } plh[idx].append(&(str.at(i)), 1); fspace = false; } for (int i = 0; i < ntok; i++) { (token + i)->SetValue(plh[i]); } // Working the tokens' IDs // Checking whether a token is a keyword, an operator or a value for (int i = 0; i < ntok; i++) { if (aKeyword((token + i)->GetValue()) == true) { (token + i)->SetID(kwd); } else if (aOperator((token + i)->GetValue()) == true) { // <--- The problem (token + i)->SetID(opr); } else { (token + i)->SetID(val); } } } // init.cpp #include "cls.h" int main() { std::string usinput; Token token[256]; std::cout << "User input: "; usinput = UserInput(); // Taking user's input Lexical(token, 256, usinput, (int) usinput.size()); // Calling the lexical function // Outputting the tokens' values and IDs (The IDs is in the parentheses) for (int i = 0; i < 10; i++) { std::cout << "Token[" << i << "] = " << token[i].GetValue() << std::setfill(' ') << std::setw(15 - (token[i].GetValue()).size()) << "(" << token[i].GetID() << ")" << std::endl; } return 0; }
问题出在Lexical函数对token的ID分类逻辑中:在枚举定义enum id {undef, kwd, idt, opr, val};里,opr(运算符)的枚举值应为3,但实际运行时,+、-等运算符的ID被错误识别为1(kwd,关键字),程序输出如下:
benanthony@DESKTOP-QI9Q4LV:~/codes/CLine$ make g++ init.cpp lex.cpp input.cpp cls.cpp -Wall -Wextra -o CLine benanthony@DESKTOP-QI9Q4LV:~/codes/CLine$ ./CLine User input: Hello there + - Rand - Lett - echo yes Token[0] = Hello (4) Token[1] = there (4) Token[2] = + (1) Token[3] = - (1) Token[4] = Rand (4) Token[5] = - (1) Token[6] = Lett (4) Token[7] = - (1) Token[8] = echo (1) Token[9] = yes (4)
我尝试过修改ID名称等方法,但问题仍未解决。请问代码中是否存在错误,或是我遗漏了关键细节?
问题分析与解决
核心错误原因
问题出在aKeyword和aOperator函数的循环边界上:
- 你定义的
keywords数组只有2个元素("echo"、"exit"),但aKeyword循环了10次(i < 10) operators数组只有4个元素("+"、"-"、"/"、"%"),aOperator同样循环了10次
当循环访问数组下标超过实际元素数量的位置时,会读取到内存中的垃圾值。这些垃圾值恰好和运算符(比如"+"、"-")匹配,导致aKeyword错误返回true,进而把运算符标记为kwd(枚举值1)。
修正方案
固定数组长度,避免越界访问:
可以给数组定义添加长度常量,或者使用std::size()获取实际长度(C++11及以上支持)。修改
aKeyword和aOperator函数:// To check whether a token is a keyword static bool aKeyword(std::string str) { const int keywordCount = std::size(keywords); for (int i = 0; i < keywordCount; i++) { if (keywords[i] == str) return true; } return false; } // To check whether a token is an operator static bool aOperator(std::string str) { const int operatorCount = std::size(operators); for (int i = 0; i < operatorCount; i++) { if (operators[i] == str) return true; } return false; }额外优化:
- 可以把
keywords和operators改成std::unordered_set,查找效率更高,同时避免数组越界问题:// 在cls.cpp中替换原数组定义 #include <unordered_set> std::unordered_set<std::string> keywords = {"echo", "exit"}; std::unordered_set<std::string> operators = {"+", "-", "/", "%"}; // 修改aKeyword和aOperator static bool aKeyword(std::string str) { return keywords.count(str) > 0; } static bool aOperator(std::string str) { return operators.count(str) > 0; } - 同时注意更新
cls.h中的声明为extern std::unordered_set<std::string> keywords;和extern std::unordered_set<std::string> operators;
- 可以把
其他潜在问题
Lexical函数中,你传入ntok=256,然后给所有256个Token赋值,但实际分割出的token数量远少于256,后续未使用的token会保留空字符串,可能导致不必要的判断。可以在分割token时记录实际数量,只处理有效token。
内容的提问来源于stack exchange,提问作者Ben Anthony
相关产品推荐
相关产品推荐

