You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Nand2Tetris汇编器开发:PLY.Lex换行符处理异常求助

问题

我正在开发Nand2Tetris汇编器,但使用Python的ply.lex库时,汇编指令总是被错误拆分。最初认为是换行符处理不当,因为后续处理时指令中总是带有"\n",尝试了多种方法均未解决:

  • 编写专门处理换行符的函数
  • 调整ignore token
  • 通过正则表达式排除换行符
  • 从源代码文件中去除换行符
  • 增大库提供的缓冲区大小

错误输出片段:

['D=M', 'AD=D-1', '0;jmp', '**MD=1\\nA', 'MD=!D',** 'MD=A-1', 'AMD=D&A', 'MD=D|A', 'AMD=M-d', 'D=m+d', 'MD=-**D\\nD', ';JGT',** 'd=m']

说明:加粗片段为指令重叠部分。

完整源代码:

from ply import lex
import termcolor
import re

Ainst_address = re.compile(r'[0-9]+')
def if_address(Ainstruct):
      return bool(Ainst_address.match(Ainstruct))

tokens = (
  "LABEL",
  "ADDRESSES",
  "CINSTRUCTION",
  "AINSTRUCTION",
  "COMMENTS"
)

t_LABEL = r'\((?:[A-Za-z0-9]+\s*(?:[A-Za-z0-9]+)?)*\)'
t_ADDRESSES = r'\@[0-9]+'
t_AINSTRUCTION = r'\@(?:[A-Za-z0-9]+)'
t_CINSTRUCTION = r'[A-Za-z0]{1,3}\s{0,2}(?:[=|;])\s{0,2}(?:(?:!|-)?){0,2}(?:[A-Za-z0-9]{1,4})\s{0,2}(?:(?:\||\+|\-|&)?(?:[A-Za-z0-9])?)|;(?:[A-Za-z0-9]{1,3})? '
t_ignore = '\t\r\f\v '  # Ignore whitespace
t_COMMENTS = r'//.*'

lex.lex(debug=0, optimize=False, reflags=re.DOTALL)

def t_NEWLINE(t):
    r'\n+'
    t.lexer.lineno += len(t.value)
    pass
    

def t_error(t):
  print(f"Illegal character '{t.value[0]}' at line {t.lexer.lineno}")
  t.lexer.skip(1)
  

def parse_source_code(source_code):
  lexer = lex.lex()
  lexer.input(source_code)
  labels = []
  Ainstr_Address = []
  Ainst_var =[]
  Cinstruction = []
  N_Cinstruction = {
        "dest": "dest" ,
        "cmp" : "cmp" ,
        "jmp" : "jmp" ,
      }
  while True:
    tok = lexer.token()
    if not tok:
      break
    if tok.type == "LABEL":
      labels.append(tok.value[1:-1])  # Remove Parenthesis from label
      print(tok.value)
    elif tok.type == "AINSTRUCTION":
        Ainstruct = (tok.value.strip("@"))
        print(termcolor.colored("Ains", "green"))
        print(tok.value)
        if  if_address(Ainstruct):
           print(termcolor.colored("Address", "blue"))
           Ainstr_Address.append(tok.value.strip("@"))
        else:
           print(f"the command is {tok.value}")
           Ainst_var.append(tok.value.strip("@"))
    elif tok.type == "CINSTRUCTION":
      print(termcolor.colored("CINSTRUCTION", "magenta"))
      Cinstruction.append(tok.value.strip("\n"))
      print(tok.value)
    elif tok.type == "COMMENTS":
      print(tok.value)
    else:
      continue

  return labels, Ainst_var, Ainstr_Address, Cinstruction
   


source_code = """
//n=2
@2
D=M
(loop)
AD=D-1
0;jmp
@ali
MD=1
AMD=!D
@moham
MD=A-1
AMD=D&A
MD=D|A
AMD=M-d
@15
D=m+d
MD=-D
D;JGT
@R1     //using a label
d=m
(example description to the developer)
"""



#total labels 3, total vars 3, total Ainst_addre 2, Cinstr 13
LAB, VARS, ADDRESSES, CINSTR = parse_source_code(source_code)
print(LAB)
print(VARS)
print(ADDRESSES)
print(CINSTR)
print(len(CINSTR))
解决方案

问题根源分析

  1. C指令正则表达式不精准:原正则末尾包含多余空格,且允许匹配不完整的指令结构,导致lexer误将换行符后的字符合并到前一条指令中。
  2. Token匹配顺序冲突:PLY按token定义的顺序进行最长匹配,原代码中AINSTRUCTION和ADDRESSES的顺序可能导致数字地址被错误识别为变量。
  3. 错误使用re.DOTALL:该flag会让正则中的.匹配换行符,导致注释和指令错误跨行匹配。
  4. 重复初始化lexer:每次调用parse_source_code都重新生成lexer实例,可能导致配置冲突。

修改后的完整代码

from ply import lex
import termcolor
import re

Ainst_address = re.compile(r'[0-9]+')
def if_address(Ainstruct):
      return bool(Ainst_address.match(Ainstruct))

# 调整token顺序,更具体的token放在前面
tokens = (
  "LABEL",
  "ADDRESSES",
  "AINSTRUCTION",
  "CINSTRUCTION",
  "COMMENTS"
)

# 修正LABEL正则,符合Nand2Tetris标签规则
t_LABEL = r'\([A-Za-z0-9_]+\)'
t_ADDRESSES = r'\@[0-9]+'
t_AINSTRUCTION = r'\@[A-Za-z0-9_]+'

# 精准匹配C指令的两种合法格式:dest=comp 或 comp;jmp
t_CINSTRUCTION = r'[ADM]{0,3}=[!\-]?[ADM01]([+\-|&][ADM01])?|[!\-]?[ADM01]([+\-|&][ADM01])?;[J][EGTLNMP]+'

t_ignore = '\t\r\f\v '  # Ignore whitespace
t_COMMENTS = r'//.*'

# 移除re.DOTALL,避免跨行匹配
lexer = lex.lex(debug=0, optimize=False)

def t_NEWLINE(t):
    r'\n+'
    t.lexer.lineno += len(t.value)
    pass
    

def t_error(t):
  print(f"Illegal character '{t.value[0]}' at line {t.lexer.lineno}")
  t.lexer.skip(1)
  

def parse_source_code(source_code):
  # 复用全局lexer实例,避免重复初始化
  lexer.input(source_code)
  labels = []
  Ainstr_Address = []
  Ainst_var =[]
  Cinstruction = []
  while True:
    tok = lexer.token()
    if not tok:
      break
    if tok.type == "LABEL":
      labels.append(tok.value[1:-1])
      print(tok.value)
    elif tok.type == "ADDRESSES":
        print(termcolor.colored("Address", "blue"))
        Ainstr_Address.append(tok.value.strip("@"))
    elif tok.type == "AINSTRUCTION":
        Ainstruct = tok.value.strip("@")
        print(termcolor.colored("Ains", "green"))
        print(tok.value)
        Ainst_var.append(Ainstruct)
    elif tok.type == "CINSTRUCTION":
      print(termcolor.colored("CINSTRUCTION", "magenta"))
      Cinstruction.append(tok.value.strip())
      print(tok.value)
    elif tok.type == "COMMENTS":
      print(tok.value)
    else:
      continue

  return labels, Ainst_var, Ainstr_Address, Cinstruction
   


source_code = """
//n=2
@2
D=M
(loop)
AD=D-1
0;jmp
@ali
MD=1
AMD=!D
@moham
MD=A-1
AMD=D&A
MD=D|A
AMD=M-d
@15
D=m+d
MD=-D
D;JGT
@R1     //using a label
d=m
(example description to the developer)
"""



#total labels 3, total vars 3, total Ainst_addre 2, Cinstr 13
LAB, VARS, ADDRESSES, CINSTR = parse_source_code(source_code)
print(LAB)
print(VARS)
print(ADDRESSES)
print(CINSTR)
print(len(CINSTR))

关键修改点说明

  1. 修正C指令正则:
    • 拆分两种合法C指令格式,严格匹配Nand2Tetris的指令规则,避免误匹配换行符或多余字符。
    • 限制dest只能是A、D、M的组合,comp只能是合法运算表达式,jmp只能是标准跳转指令。
  2. 调整Token顺序:
    • 将ADDRESSES放在AINSTRUCTION前面,确保数字地址优先被匹配,避免被错误识别为变量。
  3. 移除re.DOTALL:
    • 防止注释和指令跨行匹配,确保每行内容独立解析。
  4. 复用lexer实例:
    • 全局只初始化一次lexer,避免重复配置导致的异常。
  5. 优化指令处理逻辑:
    • 将数字地址的判断整合到Token类型中,无需额外函数判断,简化代码流程。

内容的提问来源于stack exchange,提问作者ali eltaib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:35:55