如何用Python的NLTK将词法分析中的标识符替换为Id编号
问题:将标识符替换为带序号的Id格式
我正在完成一项作业,要开发带词法分析功能的记事本程序——在记事本里写代码后,用词法分析器分词分类,最后把分类为“标识符”的标记替换成Id+序号的格式,再输出修改后的代码。现在大部分功能都完成了,唯独替换这一步卡壳了。
以下是我的代码:
def cmdAnalyze (): Analyze_program = notepad.get(0.0, END) Analyze_program_tokens = nltk.wordpunct_tokenize(Analyze_program); RE_keywords = "auto|break|case|char|const|continue|default|print" RE_Operators = "(\++)|(-)|(=)|(\*)|(/)|(%)|(--)|(<=)|(>=)" RE_Numerals = "^(\d+)$" RE_Especial_Character = "[\[@&!#$\^\|{}]:;<>?,\.']|\(\)|\(|\)|{}|\[\]|\"" RE_Identificadores = "^[a-zA-Z_]+[a-zA-Z0-9_]*" RE_Headers = "([a-zA-Z]+\.[h])" # Categorización de tokens notepad.insert(END, "\n ") for token in Analyze_program_tokens: if (re.findall(RE_keywords, token)): notepad.insert(END, "\n " + token + " --------> Palabra clave") elif (re.findall(RE_Operators, token)): notepad.insert(END, "\n " + token + " --------> Operador") elif (re.findall(RE_Numerals, token)): notepad.insert(END, "\n " + token + " --------> Número") elif (re.findall(RE_Especial_Character, token)): notepad.insert(END, "\n " + token + " --------> Carácter especial/Símbolo") elif (re.findall(RE_Identificadores, token)): notepad.insert(END, "\n " + token + " --------> Identificadores") elif (re.findall(RE_Headers, token)): notepad.insert(END, "\n " + token + " --------> Headers") else: notepad.insert(END, "\n " + " Valor desconocido") notepad.insert(END, "\n ") notepad.insert(END, Analyze_program_tokens)
当前输出:
>>> print(‘Hello World’) >>> --------> Carácter especial/Símbolo print --------> Palabra clave (‘ --------> Carácter especial/Símbolo Hello --------> Identificadores World --------> Identificadores ’) --------> Carácter especial/Símbolo >>> print (‘ Hello World ’)
我需要最后一行输出为:>>> print (‘ Id1 Id2 ’)
解决方案
要实现标识符的替换,你需要新增标识符映射记录和修改后的token列表,具体调整如下:
修改后的完整代码
import re import nltk from tkinter import END # 假设notepad是tkinter的Text组件 def cmdAnalyze (): Analyze_program = notepad.get(0.0, END).strip() # 去除首尾空白,避免多余空token Analyze_program_tokens = nltk.wordpunct_tokenize(Analyze_program) RE_keywords = "auto|break|case|char|const|continue|default|print" RE_Operators = r"(\++)|(-)|(=)|(\*)|(/)|(%)|(--)|(<=)|(>=)" # 修正转义,用原始字符串更安全 RE_Numerals = r"^(\d+)$" RE_Especial_Character = r"[\[@&!#$\^|{}]:;<>?,.']|\(|\)|{}|\[\]|\"" # 简化正则,修正转义 RE_Identificadores = r"^[a-zA-Z_]+[a-zA-Z0-9_]*$" # 加上$确保完全匹配 RE_Headers = r"([a-zA-Z]+\.[h])" # 新增:记录标识符与Id的映射,以及计数器 identifier_map = {} current_id = 1 modified_tokens = [] notepad.insert(END, "\n ") for token in Analyze_program_tokens: if re.fullmatch(RE_keywords, token): # 用fullmatch替代findall,确保完全匹配 notepad.insert(END, "\n " + token + " --------> 关键字") modified_tokens.append(token) elif re.fullmatch(RE_Operators, token): notepad.insert(END, "\n " + token + " --------> 运算符") modified_tokens.append(token) elif re.fullmatch(RE_Numerals, token): notepad.insert(END, "\n " + token + " --------> 数字") modified_tokens.append(token) elif re.fullmatch(RE_Especial_Character, token): notepad.insert(END, "\n " + token + " --------> 特殊字符/符号") modified_tokens.append(token) elif re.fullmatch(RE_Identificadores, token): notepad.insert(END, "\n " + token + " --------> 标识符") # 处理标识符替换:同一个标识符对应同一个Id if token not in identifier_map: identifier_map[token] = f"Id{current_id}" current_id += 1 modified_tokens.append(identifier_map[token]) elif re.fullmatch(RE_Headers, token): notepad.insert(END, "\n " + token + " --------> 头文件") modified_tokens.append(token) else: notepad.insert(END, "\n " + " 未知值") modified_tokens.append(token) # 将修改后的token拼接成字符串并插入 modified_code = ' '.join(modified_tokens) notepad.insert(END, "\n ") notepad.insert(END, modified_code)
关键调整说明
- 标识符映射机制:用
identifier_map字典记录每个标识符对应的Id,确保重复出现的标识符会被替换为同一个Id(如果不需要去重,直接每次计数递增即可,去掉字典判断)。 - 正则匹配优化:把
re.findall换成re.fullmatch,避免部分匹配导致的分类错误;同时将正则改为原始字符串(r""),减少转义错误。 - 生成修改后的代码:用
modified_tokens列表收集替换后的所有token,最后用' '.join()拼接成完整代码插入记事本。 - 处理空白:读取文本时用
.strip()去除首尾空白,避免生成多余的空token。
内容的提问来源于stack exchange,提问作者Jhonyjosaku
相关产品推荐
相关产品推荐

