基于编译原理逐字符解析字符串时如何忽略//和{}格式的注释
解决方案
实现思路
- 新增两个状态标记分别记录当前是否处于单行注释、块注释状态
- 遍历字符时优先判断注释状态,处于注释状态时跳过字符不加入临时缓存
- 两类注释的触发/终止逻辑:
- 单行注释:连续遇到
/和/触发,遇到换行符终止 - 块注释:遇到
{触发,遇到}终止
- 单行注释:连续遇到
完整修改代码
string = """ beGIn west WEST north//comment1 \n north north west East east south\n // comment west\n {\n comment\n }\n end """ tokens = [] tmp = '' in_line_comment = False in_block_comment = False for i, peek in enumerate(string.lower()): # 优先处理块注释状态 if in_block_comment: if peek == '}': in_block_comment = False continue # 处理单行注释状态 if in_line_comment: if peek == '\n': in_line_comment = False continue # 判断是否进入单行注释 if peek == '/' and i > 0 and string[i-1].lower() == '/': # 删掉之前多加入的单个/ tmp = tmp.rstrip('/') in_line_comment = True continue # 判断是否进入块注释 if peek == '{': in_block_comment = True continue # 原有空白符处理逻辑 if peek == ' ' or peek == '\n': if len(tmp) > 0: tokens.append(tmp) print(tmp) tmp = '' else: tmp += peek # 处理末尾未被空白符触发的剩余token if len(tmp) > 0: tokens.append(tmp) print(tmp)
输出结果
运行上述代码后输出完全符合预期:
begin west west north north north west east east south end
内容的提问来源于stack exchange,提问作者sep_The_new_elixiR
相关产品推荐
相关产品推荐

