You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的NLTK将词法分析中的标识符替换为Id编号

问题:将标识符替换为带序号的Id格式

我正在完成一项作业,要开发带词法分析功能的记事本程序——在记事本里写代码后,用词法分析器分词分类,最后把分类为“标识符”的标记替换成Id+序号的格式,再输出修改后的代码。现在大部分功能都完成了,唯独替换这一步卡壳了。

以下是我的代码:

def cmdAnalyze (): 

    Analyze_program = notepad.get(0.0, END)
    Analyze_program_tokens = nltk.wordpunct_tokenize(Analyze_program);

    RE_keywords = "auto|break|case|char|const|continue|default|print"
    RE_Operators = "(\++)|(-)|(=)|(\*)|(/)|(%)|(--)|(<=)|(>=)"
    RE_Numerals = "^(\d+)$"
    RE_Especial_Character = "[\[@&!#$\^\|{}]:;<>?,\.']|\(\)|\(|\)|{}|\[\]|\""
    RE_Identificadores = "^[a-zA-Z_]+[a-zA-Z0-9_]*"
    RE_Headers = "([a-zA-Z]+\.[h])"

    # Categorización de tokens

    notepad.insert(END, "\n ")

    for token in Analyze_program_tokens:
        if (re.findall(RE_keywords, token)):
            notepad.insert(END, "\n " + token + " --------> Palabra clave")
        elif (re.findall(RE_Operators, token)):
            notepad.insert(END, "\n " + token + " --------> Operador")
        elif (re.findall(RE_Numerals, token)):
            notepad.insert(END, "\n " + token + " --------> Número")
        elif (re.findall(RE_Especial_Character, token)):
            notepad.insert(END, "\n " + token + " --------> Carácter especial/Símbolo")
        elif (re.findall(RE_Identificadores, token)):
            notepad.insert(END, "\n " + token + " --------> Identificadores")
        elif (re.findall(RE_Headers, token)):
            notepad.insert(END, "\n " + token + " --------> Headers")

        else:
            notepad.insert(END, "\n " + " Valor desconocido")

    notepad.insert(END, "\n ")
    notepad.insert(END, Analyze_program_tokens)

当前输出:

>>> print(‘Hello World’)

 >>> --------> Carácter especial/Símbolo
 print --------> Palabra clave
 (‘ --------> Carácter especial/Símbolo
 Hello --------> Identificadores
 World --------> Identificadores
 ’) --------> Carácter especial/Símbolo
 >>> print (‘ Hello World ’)

我需要最后一行输出为:>>> print (‘ Id1 Id2 ’)


解决方案

要实现标识符的替换,你需要新增标识符映射记录和修改后的token列表,具体调整如下:

修改后的完整代码

import re
import nltk
from tkinter import END  # 假设notepad是tkinter的Text组件

def cmdAnalyze (): 
    Analyze_program = notepad.get(0.0, END).strip()  # 去除首尾空白,避免多余空token
    Analyze_program_tokens = nltk.wordpunct_tokenize(Analyze_program)

    RE_keywords = "auto|break|case|char|const|continue|default|print"
    RE_Operators = r"(\++)|(-)|(=)|(\*)|(/)|(%)|(--)|(<=)|(>=)"  # 修正转义,用原始字符串更安全
    RE_Numerals = r"^(\d+)$"
    RE_Especial_Character = r"[\[@&!#$\^|{}]:;<>?,.']|\(|\)|{}|\[\]|\""  # 简化正则,修正转义
    RE_Identificadores = r"^[a-zA-Z_]+[a-zA-Z0-9_]*$"  # 加上$确保完全匹配
    RE_Headers = r"([a-zA-Z]+\.[h])"

    # 新增:记录标识符与Id的映射,以及计数器
    identifier_map = {}
    current_id = 1
    modified_tokens = []

    notepad.insert(END, "\n ")

    for token in Analyze_program_tokens:
        if re.fullmatch(RE_keywords, token):  # 用fullmatch替代findall,确保完全匹配
            notepad.insert(END, "\n " + token + " --------> 关键字")
            modified_tokens.append(token)
        elif re.fullmatch(RE_Operators, token):
            notepad.insert(END, "\n " + token + " --------> 运算符")
            modified_tokens.append(token)
        elif re.fullmatch(RE_Numerals, token):
            notepad.insert(END, "\n " + token + " --------> 数字")
            modified_tokens.append(token)
        elif re.fullmatch(RE_Especial_Character, token):
            notepad.insert(END, "\n " + token + " --------> 特殊字符/符号")
            modified_tokens.append(token)
        elif re.fullmatch(RE_Identificadores, token):
            notepad.insert(END, "\n " + token + " --------> 标识符")
            # 处理标识符替换:同一个标识符对应同一个Id
            if token not in identifier_map:
                identifier_map[token] = f"Id{current_id}"
                current_id += 1
            modified_tokens.append(identifier_map[token])
        elif re.fullmatch(RE_Headers, token):
            notepad.insert(END, "\n " + token + " --------> 头文件")
            modified_tokens.append(token)
        else:
            notepad.insert(END, "\n " + " 未知值")
            modified_tokens.append(token)

    # 将修改后的token拼接成字符串并插入
    modified_code = ' '.join(modified_tokens)
    notepad.insert(END, "\n ")
    notepad.insert(END, modified_code)

关键调整说明

  1. 标识符映射机制:用identifier_map字典记录每个标识符对应的Id,确保重复出现的标识符会被替换为同一个Id(如果不需要去重,直接每次计数递增即可,去掉字典判断)。
  2. 正则匹配优化:把re.findall换成re.fullmatch,避免部分匹配导致的分类错误;同时将正则改为原始字符串(r""),减少转义错误。
  3. 生成修改后的代码:用modified_tokens列表收集替换后的所有token,最后用' '.join()拼接成完整代码插入记事本。
  4. 处理空白:读取文本时用.strip()去除首尾空白,避免生成多余的空token。

内容的提问来源于stack exchange,提问作者Jhonyjosaku

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 19:05:28