如何判断Python单行代码是否含注释并拆分出code与comment两部分
功能需求
- 判断单行Python代码是否包含注释,返回True/False结果
- 将该行代码拆分为
code(代码部分)和comment(注释部分)两个部分
场景说明
普通带注释的代码处理逻辑非常简单,示例如下:
loc_1 = "print('hello') # this is a comment"
复杂场景下需要识别字符串内部的#不属于注释,不能作为拆分边界,示例如下:
loc_2 = 'for char in "(*#& eht # ": # pylint: disable=one,two # something'
预期输出:
调用处理函数后返回结果如下:
f(loc_2) # returns [ # code 'for char in "(*#& eht # ":', # comment ' # pylint: disable=one,two # something' ]
背景说明
- 曾尝试使用libcst库获取AST实现需求,未成功
- 处理对象为单行代码:现有流程会逐行遍历Python文件,需要扩展注释识别功能
- 最终目标是分别得到完整的代码部分和注释部分内容,而非仅获取无注释源码
实现方案
使用Python标准库tokenize实现,无需依赖第三方库,可以精准区分字符串内的#和注释起始的#,完全满足需求:
import tokenize from io import BytesIO def split_code_and_comment(line: str) -> tuple[bool, str, str]: """ 拆分单行Python代码为代码部分和注释部分 返回值:(是否包含注释, 代码部分, 注释部分) """ line_bytes = line.encode('utf-8') tokens = list(tokenize.tokenize(BytesIO(line_bytes).readline)) comment_start = len(line) has_comment = False for tok in tokens: tok_type, tok_string, (_, start_col), _, _ = tok if tok_type == tokenize.COMMENT: comment_start = start_col has_comment = True break code_part = line[:comment_start].rstrip() comment_part = line[comment_start:] return has_comment, code_part, comment_part # 测试示例 if __name__ == "__main__": loc2 = 'for char in "(*#& eht # ": # pylint: disable=one,two # something' has_comment, code, comment = split_code_and_comment(loc2) print(has_comment) # 输出 True print(repr(code)) # 输出 'for char in "(*#& eht # ":' print(repr(comment)) # 输出 ' # pylint: disable=one,two # something'
实现原理:tokenize是Python官方提供的源码分词工具,会自动识别字符串、注释、关键字等不同语法单元,碰到COMMENT类型的token时,对应的起始列就是注释的开始位置,按位置拆分即可,不会出现字符串内#的误判问题。
内容的提问来源于stack exchange,提问作者baxx
相关产品推荐
相关产品推荐

