Nand2Tetris汇编器开发:PLY.Lex换行符处理异常求助
问题
我正在开发Nand2Tetris汇编器,但使用Python的ply.lex库时,汇编指令总是被错误拆分。最初认为是换行符处理不当,因为后续处理时指令中总是带有"\n",尝试了多种方法均未解决:
- 编写专门处理换行符的函数
- 调整ignore token
- 通过正则表达式排除换行符
- 从源代码文件中去除换行符
- 增大库提供的缓冲区大小
错误输出片段:
['D=M', 'AD=D-1', '0;jmp', '**MD=1\\nA', 'MD=!D',** 'MD=A-1', 'AMD=D&A', 'MD=D|A', 'AMD=M-d', 'D=m+d', 'MD=-**D\\nD', ';JGT',** 'd=m']
说明:加粗片段为指令重叠部分。
完整源代码:
from ply import lex import termcolor import re Ainst_address = re.compile(r'[0-9]+') def if_address(Ainstruct): return bool(Ainst_address.match(Ainstruct)) tokens = ( "LABEL", "ADDRESSES", "CINSTRUCTION", "AINSTRUCTION", "COMMENTS" ) t_LABEL = r'\((?:[A-Za-z0-9]+\s*(?:[A-Za-z0-9]+)?)*\)' t_ADDRESSES = r'\@[0-9]+' t_AINSTRUCTION = r'\@(?:[A-Za-z0-9]+)' t_CINSTRUCTION = r'[A-Za-z0]{1,3}\s{0,2}(?:[=|;])\s{0,2}(?:(?:!|-)?){0,2}(?:[A-Za-z0-9]{1,4})\s{0,2}(?:(?:\||\+|\-|&)?(?:[A-Za-z0-9])?)|;(?:[A-Za-z0-9]{1,3})? ' t_ignore = '\t\r\f\v ' # Ignore whitespace t_COMMENTS = r'//.*' lex.lex(debug=0, optimize=False, reflags=re.DOTALL) def t_NEWLINE(t): r'\n+' t.lexer.lineno += len(t.value) pass def t_error(t): print(f"Illegal character '{t.value[0]}' at line {t.lexer.lineno}") t.lexer.skip(1) def parse_source_code(source_code): lexer = lex.lex() lexer.input(source_code) labels = [] Ainstr_Address = [] Ainst_var =[] Cinstruction = [] N_Cinstruction = { "dest": "dest" , "cmp" : "cmp" , "jmp" : "jmp" , } while True: tok = lexer.token() if not tok: break if tok.type == "LABEL": labels.append(tok.value[1:-1]) # Remove Parenthesis from label print(tok.value) elif tok.type == "AINSTRUCTION": Ainstruct = (tok.value.strip("@")) print(termcolor.colored("Ains", "green")) print(tok.value) if if_address(Ainstruct): print(termcolor.colored("Address", "blue")) Ainstr_Address.append(tok.value.strip("@")) else: print(f"the command is {tok.value}") Ainst_var.append(tok.value.strip("@")) elif tok.type == "CINSTRUCTION": print(termcolor.colored("CINSTRUCTION", "magenta")) Cinstruction.append(tok.value.strip("\n")) print(tok.value) elif tok.type == "COMMENTS": print(tok.value) else: continue return labels, Ainst_var, Ainstr_Address, Cinstruction source_code = """ //n=2 @2 D=M (loop) AD=D-1 0;jmp @ali MD=1 AMD=!D @moham MD=A-1 AMD=D&A MD=D|A AMD=M-d @15 D=m+d MD=-D D;JGT @R1 //using a label d=m (example description to the developer) """ #total labels 3, total vars 3, total Ainst_addre 2, Cinstr 13 LAB, VARS, ADDRESSES, CINSTR = parse_source_code(source_code) print(LAB) print(VARS) print(ADDRESSES) print(CINSTR) print(len(CINSTR))
解决方案
问题根源分析
- C指令正则表达式不精准:原正则末尾包含多余空格,且允许匹配不完整的指令结构,导致lexer误将换行符后的字符合并到前一条指令中。
- Token匹配顺序冲突:PLY按token定义的顺序进行最长匹配,原代码中
AINSTRUCTION和ADDRESSES的顺序可能导致数字地址被错误识别为变量。 - 错误使用
re.DOTALL:该flag会让正则中的.匹配换行符,导致注释和指令错误跨行匹配。 - 重复初始化lexer:每次调用
parse_source_code都重新生成lexer实例,可能导致配置冲突。
修改后的完整代码
from ply import lex import termcolor import re Ainst_address = re.compile(r'[0-9]+') def if_address(Ainstruct): return bool(Ainst_address.match(Ainstruct)) # 调整token顺序,更具体的token放在前面 tokens = ( "LABEL", "ADDRESSES", "AINSTRUCTION", "CINSTRUCTION", "COMMENTS" ) # 修正LABEL正则,符合Nand2Tetris标签规则 t_LABEL = r'\([A-Za-z0-9_]+\)' t_ADDRESSES = r'\@[0-9]+' t_AINSTRUCTION = r'\@[A-Za-z0-9_]+' # 精准匹配C指令的两种合法格式:dest=comp 或 comp;jmp t_CINSTRUCTION = r'[ADM]{0,3}=[!\-]?[ADM01]([+\-|&][ADM01])?|[!\-]?[ADM01]([+\-|&][ADM01])?;[J][EGTLNMP]+' t_ignore = '\t\r\f\v ' # Ignore whitespace t_COMMENTS = r'//.*' # 移除re.DOTALL,避免跨行匹配 lexer = lex.lex(debug=0, optimize=False) def t_NEWLINE(t): r'\n+' t.lexer.lineno += len(t.value) pass def t_error(t): print(f"Illegal character '{t.value[0]}' at line {t.lexer.lineno}") t.lexer.skip(1) def parse_source_code(source_code): # 复用全局lexer实例,避免重复初始化 lexer.input(source_code) labels = [] Ainstr_Address = [] Ainst_var =[] Cinstruction = [] while True: tok = lexer.token() if not tok: break if tok.type == "LABEL": labels.append(tok.value[1:-1]) print(tok.value) elif tok.type == "ADDRESSES": print(termcolor.colored("Address", "blue")) Ainstr_Address.append(tok.value.strip("@")) elif tok.type == "AINSTRUCTION": Ainstruct = tok.value.strip("@") print(termcolor.colored("Ains", "green")) print(tok.value) Ainst_var.append(Ainstruct) elif tok.type == "CINSTRUCTION": print(termcolor.colored("CINSTRUCTION", "magenta")) Cinstruction.append(tok.value.strip()) print(tok.value) elif tok.type == "COMMENTS": print(tok.value) else: continue return labels, Ainst_var, Ainstr_Address, Cinstruction source_code = """ //n=2 @2 D=M (loop) AD=D-1 0;jmp @ali MD=1 AMD=!D @moham MD=A-1 AMD=D&A MD=D|A AMD=M-d @15 D=m+d MD=-D D;JGT @R1 //using a label d=m (example description to the developer) """ #total labels 3, total vars 3, total Ainst_addre 2, Cinstr 13 LAB, VARS, ADDRESSES, CINSTR = parse_source_code(source_code) print(LAB) print(VARS) print(ADDRESSES) print(CINSTR) print(len(CINSTR))
关键修改点说明
- 修正C指令正则:
- 拆分两种合法C指令格式,严格匹配Nand2Tetris的指令规则,避免误匹配换行符或多余字符。
- 限制dest只能是A、D、M的组合,comp只能是合法运算表达式,jmp只能是标准跳转指令。
- 调整Token顺序:
- 将
ADDRESSES放在AINSTRUCTION前面,确保数字地址优先被匹配,避免被错误识别为变量。
- 将
- 移除
re.DOTALL:- 防止注释和指令跨行匹配,确保每行内容独立解析。
- 复用lexer实例:
- 全局只初始化一次lexer,避免重复配置导致的异常。
- 优化指令处理逻辑:
- 将数字地址的判断整合到Token类型中,无需额外函数判断,简化代码流程。
内容的提问来源于stack exchange,提问作者ali eltaib
相关产品推荐
相关产品推荐

