如何对存在语法错误的Python源代码进行分词处理?
处理含语法错误的Python代码分词问题
问题背景
我需要对存在语法错误的Python源代码进行分词,用于输入到循环神经网络(RNN)这类统计模型中。但Python内置的tokenize模块在处理这类代码时,要么生成ErrorToken,要么直接抛出异常,无法完成完整分词。
现有实现代码
我目前使用的分词函数如下:
from typing import List from io import BytesIO import tokenize def to_token_list(s: str) -> List: tokens = [] # 从源代码提取的token列表 g = tokenize.tokenize(BytesIO(s.encode("utf-8")).readline) for t in g: tokens.append(t) return tokens
问题复现示例
示例1:缺少闭合括号的代码
输入代码(对象已掩码为ID):
syntax_error_source_code = "\ndef ID ID ):\n if ID .ID :\n ID .ID .ID ()\n" to_token_list(syntax_error_source_code)
触发报错:
Exception has occurred: TokenError ('EOF in multi-line statement', (5, 0))
示例2:缩进错误的代码
输入代码:
syntax_error_source_code = '\ndef ID ():\n/ for ID ,ID in ID :\n pass \n for ID ,ID in ID :\n pass \n' to_token_list(syntax_error_source_code)
触发报错:
Exception has occurred: IndentationError unindent does not match any outer indentation level (<tokenize>, line 5)
尝试过的方法
我试过用try-except捕获异常,但这种方式只能处理末尾的错误,对于代码中间出现的错误仍然无法解决,无法完成完整的分词流程。
寻求解决方案
请问有没有办法规避这类问题,实现对含语法错误的Python代码的完整分词?
内容的提问来源于stack exchange,提问作者semyd
相关产品推荐
相关产品推荐

