You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对存在语法错误的Python源代码进行分词处理?

处理含语法错误的Python代码分词问题

问题背景

我需要对存在语法错误的Python源代码进行分词,用于输入到循环神经网络(RNN)这类统计模型中。但Python内置的tokenize模块在处理这类代码时,要么生成ErrorToken,要么直接抛出异常,无法完成完整分词。

现有实现代码

我目前使用的分词函数如下:

from typing import List
from io import BytesIO
import tokenize

def to_token_list(s: str) -> List:
    tokens = []  # 从源代码提取的token列表

    g = tokenize.tokenize(BytesIO(s.encode("utf-8")).readline)

    for t in g:
        tokens.append(t)

    return tokens

问题复现示例

示例1:缺少闭合括号的代码

输入代码(对象已掩码为ID):

syntax_error_source_code = "\ndef ID ID ):\n    if ID .ID :\n        ID .ID .ID ()\n"
to_token_list(syntax_error_source_code)

触发报错:

Exception has occurred: TokenError
('EOF in multi-line statement', (5, 0))

示例2:缩进错误的代码

输入代码:

syntax_error_source_code = '\ndef ID ():\n/    for ID ,ID in ID :\n        pass \n    for ID ,ID in ID :\n        pass \n'
to_token_list(syntax_error_source_code)

触发报错:

Exception has occurred: IndentationError
unindent does not match any outer indentation level (<tokenize>, line 5)

尝试过的方法

我试过用try-except捕获异常,但这种方式只能处理末尾的错误,对于代码中间出现的错误仍然无法解决,无法完成完整的分词流程。

寻求解决方案

请问有没有办法规避这类问题,实现对含语法错误的Python代码的完整分词?

内容的提问来源于stack exchange,提问作者semyd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 07:00:57